VLDB 2026 Research / reviewers in the wild / expert
Minglun Gong
dblp:18/1338
· DBLP profile ↗
114ranked-venue papers
21as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 87 · 14 first-author · 18 since 2021Artificial intelligence and machine learning · 46 · 12 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ParkinsonNet: A unified end-to-end framework for estimating Parkinson's disease motor symptom severityabstract• ParkinsonNet: We propose a novel end-to-end network, ParkinsonNet, for automatically qualifying the severity of PD motor symptoms. This network is capable of processing multiple motor tests without any manual feature design. • Temporal Self-Attention Enhancement Module (TAEM): To accurately perceive the gradual progression of motor symptoms, we introduce TAEM, which combines temporal compression with long-term dependency modeling, enabling robust and comprehensive temporal feature extraction. • Similarity Matching Module (SMM): To address the challenges of class imbalance and limited datasets, we design the SMM, that transforms the conventional classification or regression task into a similarity matching problem. This module aligns skeleton features with their most similar text-based features, leveraging semantic relationships for improved performance. • Empirical evaluations: Extensive experiments are conducted on two recent PD datasets, demonstrating the superiority of ParkinsonNet, and providing essential benchmark performance for future algorithm development and evaluation. Parkinson’s Disease (PD) is a progressive neurodegenerative disorder characterized by worsening motor symptoms such as bradykinesia, imbalance, tremor, rigidity, and gait disturbances. Clinician assessments are often time-consuming and costly, and the limited availability of specialists, along with patient mobility issues, complicates frequent evaluations. In this paper, we propose a novel end-to-end network to automatically qualify the severity of motor symptoms in PD, referred to as ParkinsonNet. Unlike most existing methods that focus on isolated tests, ParkinsonNet provides a unified learning framework that is evaluated across multiple PD motor symptoms, as demonstrated on finger tapping and gait. Specifically, to accurately perceive the gradual progression of motor symptoms throughout an entire test cycle (e.g., decrementing amplitude), a temporal self-attention enhancement module is designed by combining temporal compression with long-term temporal dependency modeling. To ease the issues of class imbalance and limited datasets, a similarity matching module is proposed that transforms the conventional classification or regression task into a similarity matching problem, matching the skeleton feature with its most similar texture feature. Additionally, a vector quantization module is incorporated to encode spatiotemporal features into a discrete-valued space, compressing and abstracting motion representations while retaining critical information for more accurate classification. Extensive experiments on two newly identified benchmark datasets demonstrate the superiority of our ParkinsonNet and set new benchmark performance for future algorithm development and evaluation. Our code will be released at this https URL . Yande Li, Fang Ba, Minglun Gong, Li Cheng 0001 |
Pattern Recognit. | 3 |
| 2026 | VLCounting: Taming zero-shot counting via language-driven exemplar grounding
Mingjie Wang 0002, Yong Dai 0001, Eric Buys, Minglun Gong |
Pattern Recognit. | 5 |
| 2025 | ArcPro: Architectural Programs for Structured 3D Abstraction of Sparse PointsabstractWe introduce ArcPro, a novel learning framework built on architectural programs to recover structured 3D abstractions from highly sparse and low-quality point clouds. Specifically, we design a domain-specific language (DSL) to hierarchically represent building structures as a program, which can be efficiently converted into a mesh. We bridge feedforward and inverse procedural modeling by using a feedforward process for training data synthesis, allowing the network to make reverse predictions. We train an encoder-decoder on the points-program pairs to establish a mapping from unstructured point clouds to architectural programs, where a 3D convolutional encoder extracts point cloud features and a transformer decoder autoregressively predicts the programs in a tokenized form. Inference by our method is highly efficient and produces plausible and faithful 3D abstractions. Comprehensive experiments demonstrate that ArcPro outperforms both traditional architectural proxy reconstruction and learning-based abstraction methods. We further explore its potential to work with multi-view image and natural language inputs. Project page: https://vcc.tech/research/2025/ArcPro. Kangjun Liu, Minglun Gong, Hao (Richard) Zhang, Hui Huang 0004 |
CVPR | 4 |
| 2025 | Textured-GS: Gaussian Splatting with Spatially Defined Color and OpacityabstractIn this paper, we introduce Textured-GS, an innovative method for improving the rendering quality of 3D Gaussian splatting that incorporates spatially defined color and opacity variations using Spherical Harmonics (SH). This approach enables each Gaussian to exhibit a richer representation by accommodating varying colors and opacities across its surface, significantly enhancing rendering quality compared to traditional methods. To demonstrate the merits of our approach, we adapted the Mini-Splatting architecture to integrate textured Gaussians without increasing the number of Gaussians. Our experiments across multiple real-world datasets show that Textured-GS consistently outperforms both the baseline Mini-Splatting and standard 3DGS in terms of visual fidelity. These results highlight the potential of Textured-GS to advance Gaussian-based rendering technologies, promising more efficient and high-quality scene reconstructions. Our implementation is available at https://github.com/ZhentaoHuang/Textured-GS. Minglun Gong |
Graphics Interface | 2 |
| 2025 | GRIG: Data-Efficient Generative Residual Image InpaintingabstractImage inpainting is the task of filling in missing or masked regions of an image with semantically meaningful content. Recent methods have shown significant improvement in dealing with large missing regions. However, these methods usually require large training datasets to achieve satisfactory results, and there has been limited research into training such models on a small number of samples. To address this, we present a novel data-efficient generative residual image inpainting method that produces high-quality inpainting results. The core idea is to use an iterative residual reasoning method that incorporates convolutional neural networks (CNNs) for feature extraction and transformers for global reasoning within generative adversarial networks, along with image-level and patch-level discriminators. We also propose a novel forged-patch adversarial training strategy to create faithful textures and detailed appearances. Extensive evaluation shows that our method outperforms previous methods on the data-efficient image inpainting task, both quantitatively and qualitatively. Wanglong Lu, Xianta Jiang, Xiaogang Jin 0001, Minglun Gong, Kaijie Shi 0002, Tao Wang 0052, Hanli Zhao |
Comput. Vis. Media | 5 |
| 2025 | Real-time dual-eye collaborative eyeblink detection with contrastive learning
Hanli Zhao, Yu Wang 0219, Wanglong Lu, Zili Yi, Minglun Gong |
Pattern Recognit. | 6 |
| 2025 | FER-Former: Multimodal Transformer for Facial Expression RecognitionabstractThe ever-increasing demands for intuitive interactions in virtual reality have led to surging interests in facial expression recognition (FER). There are however several issues commonly seen in existing methods, including narrow receptive fields and homogenous supervisory signals. To address these issues, we propose in this paper a novel multimodal supervision-steering transformer for facial expression recognition in the wild, referred to as FER-former. Specifically, to address the limitation of narrow receptive fields, a hybrid feature extraction pipeline is designed by cascading both prevailing CNNs and transformers. To deal with the issue of homogenous supervisory signals, a heterogeneous domain-steering supervision module is proposed to incorporate text-space semantic correlations to enhance image features, based on the similarity between image and text features. Additionally, a FER-specific transformer encoder is introduced to characterize conventional one-hot label-focusing and CLIP-based text-oriented tokens in parallel for final classification. Based on the collaboration of multifarious token heads, global receptive fields with multimodal semantic cues are captured, delivering superb learning capability. Extensive experiments on popular benchmarks demonstrate the superiority of the proposed FER-former over the existing state-of-the-art methods. Yande Li, Mingjie Wang 0002, Minglun Gong, Yonggang Lu, Li Liu 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Visibility-Aware Pixelwise View Selection for Multi-View Stereo Matching
Yukun Shi, Minglun Gong |
ICPR (18) | 3 |
| 2024 | GCNet: Probing self-similarity learning for Generalized Counting Network
Mingjie Wang 0002, Yande Li, Graham W. Taylor, Minglun Gong |
Pattern Recognit. | 5 |
| 2024 | Architectural Co-LOD GenerationabstractManaging the level-of-detail (LOD) in architectural models is crucial yet challenging, particularly for effective representation and visualization of buildings. Traditional approaches often fail to deliver controllable detail alongside semantic consistency, especially when dealing with noisy and inconsistent inputs. We address these limitations with Co-LOD , a new approach specifically designed for effective LOD management in architectural modeling. Co-LOD employs shape co-analysis to standardize geometric structures across multiple buildings, facilitating the progressive and consistent generation of LODs. This method allows for precise detailing in both individual models and model collections, ensuring semantic integrity. Extensive experiments demonstrate that Co-LOD effectively applies accurate LOD across a variety of architectural inputs, consistently delivering superior detail and quality in LOD representations. Shanshan Pan, Chenlei Lv, Minglun Gong, Hui Huang 0004 |
ACM Trans. Graph. | 4 |
| 2023 | Where to Render: Studying Renderability for IBR of Large-Scale ScenesabstractImage-based rendering (IBR) technique enables presenting real scenes interactively to viewers and hence is a key component for implementing VR telepresence. The quality of IBR results depends on the set of pre-captured views, the rendering algorithm used, and the camera parameters of the novel view to be synthesized. Numerous methods were proposed for optimizing the set of captured images and enhancing the rendering algorithms. However, from which regions IBR methods can synthesize satisfactory results is not yet well studied. In this work, we introduce the concept of renderability, which predicts the quality of IBR results at any given viewpoint and view direction. Consequently, the renderability values evaluated for the 5D camera parameter space form a field, which effectively guides viewpoint/trajectory selection for IBR, especially for challenging large-scale 3D scenes. To demonstrate this capability, we designed 2 VR applications: a path planner that allows users to navigate through sparsely captured scenes with controllable rendering quality and a view selector that provides an overview for a scene from diverse and high quality perspectives. We believe the renderability concept, the proposed evaluation method, and the suggested applications will motivate and facilitate the use of IBR in various interactive settings. Zimu Yi, Ke Xie 0001, Jiahui Lyu, Minglun Gong, Hui Huang 0004 |
VR | 4 |
| 2023 | Dynamic Mixture of Counter Network for Location-Agnostic Crowd CountingabstractCrowd counting has attracted increasing attentions in recent years due to its challenges and wide societal applications. Despite persevering efforts made by the research community, most of existing methods require a large amount of location-level annotations. Collecting such type of fine-granularity supervisory signals is extremely time-consuming and labour-intensive, thereby hindering the well generalization of these location-adherent models. To shun this drawback, several pioneering studies open a promising research direction of location-agonistic crowd counting. Albeit the noticeable efforts, they somewhat ignore the merits of diverse learning paradigms and the issue of intractable density shift. To ameliorate these issues, in this paper, a novel Dynamic Mixture of Counter Network (DMCNet) is proposed for location-agnostic crowd counting. Specifically, our DMCNet inherits the hybrid advantages of CNNs (e.g. locality-oriented and pyramidal property) and MLP-based structure (e.g. global receptive fields and light weight). Particularly, the dynamic counter predictor and the mixture of counter heads are delicately designed to hammer at combating huge density shift and overfitting. Extensive experiments demonstrate that our DMCNet attains state-of-the-art performance against existing location-agnostic approaches and performs on par with many conventional location-adherent ones. Mingjie Wang 0002, Hao Cai 0004, Yong Dai 0001, Minglun Gong |
WACV | 4 |
| 2023 | Enhanced face aging using dual-learning and muti-attention mechanism
Xin Huang 0030, Minglun Gong |
Appl. Intell. | 2 |
| 2023 | PA-Net: Plane Attention Network for real-time urban scene reconstruction
Ruiqi Cui, Ke Xie 0001, Minglun Gong, Hui Huang 0004 |
Comput. Graph. | 4 |
| 2023 | CrowdMLP: Weakly-supervised crowd counting via multi-granularity MLP
Mingjie Wang 0002, Jun Zhou 0001, Hao Cai 0004, Minglun Gong |
Pattern Recognit. | 4 |
| 2023 | STNet: Scale Tree Network With Multi-Level Auxiliator for Crowd CountingabstractState-of-the-art approaches for crowd counting resort to deepneural networks to predict density maps. However, counting people in congested scenes remains a challenging task because the presence of drastic scale variation, density inconsistency, and complex background can seriously degrade their counting accuracy. To battle the ingrained issue of accuracy degradation, in this paper, we propose a novel and powerful network called Scale Tree Network (STNet) for accurate crowd counting. STNet consists of two key components: a Scale-Tree Diversity Enhancer and a Multi-level Auxiliator. Specifically, the Diversity Enhancer is designed to enrich scale diversity, which alleviates limitations of existing methods caused by insufficient level of scales. A novel tree structure is adopted to hierarchically parse coarse-to-fine crowd regions. Furthermore, a simple yet effective Multi-level Auxiliator is presented to aid in exploiting generalisable shared characteristics at multiple levels, allowing more accurate pixel-wise background cognition. The overall STNet is trained in an end-to-end manner, without the needs for manually tuning loss weights between the main and the auxiliary tasks. Extensive experiments on five challenging crowd datasets demonstrate the superiority of the proposed method. Mingjie Wang 0002, Hao Cai 0004, Xian-Feng Han, Jun Zhou 0023, Minglun Gong |
IEEE Trans. Multim. | 5 |
| 2023 | Neural Packing: from Visual Sensing to Reinforcement LearningabstractWe present a novel learning framework to solve the transport-and-packing (TAP) problem in 3D. It constitutes a full solution pipeline from partial observations of input objects via RGBD sensing and recognition to final box placement, via robotic motion planning, to arrive at a compact packing in a target container. The technical core of our method is a neural network for TAP, trained via reinforcement learning (RL), to solve the NP-hard combinatorial optimization problem. Our network simultaneously selects an object to pack and determines the final packing location, based on a judicious encoding of the continuously evolving states of partially observed source objects and available spaces in the target container, using separate encoders both enabled with attention mechanisms. The encoded feature vectors are employed to compute the matching scores and feasibility masks of different pairings of box selection and available space configuration for packing strategy optimization. Extensive experiments, including ablation studies and physical packing execution by a real robot (Universal Robot UR5e), are conducted to evaluate our method in terms of its design choices, scalability, generalizability, and comparisons to baselines, including the most recent RL-based TAP solution. We also contribute the first benchmark for TAP which covers a variety of input settings and difficulty levels. Juzhan Xu, Minglun Gong, Hao (Richard) Zhang, Hui Huang 0004, Ruizhen Hu |
ACM Trans. Graph. | 2 |
| 2022 | Object Wake-Up: 3D Object Rigging from a Single Image
Xinxin Zuo, Sen Wang 0003, Zhenbo Yu, Bingbing Ni, Minglun Gong, Li Cheng 0001 |
ECCV (2) | 7 |
| 2022 | Action2video: Generating Videos of Human 3D Actions
Chuan Guo 0002, Xinxin Zuo, Sen Wang 0003, Xinshuang Liu, Shihao Zou, Minglun Gong, Li Cheng 0001 |
Int. J. Comput. Vis. | 6 |
| 2022 | Dual-channel feature disentanglement for identity-invariant facial expression recognition
Yande Li, Yonggang Lu, Minglun Gong, Li Liu 0001, Ligang Zhao |
Inf. Sci. | 3 |
| 2022 | 3D pose estimation and future motion prediction from 2D images
Youdong Ma, Xinxin Zuo, Sen Wang 0003, Minglun Gong, Li Cheng 0001 |
Pattern Recognit. | 5 |
| 2022 | A robust framework for multi-view stereopsis
Wendong Mao, Mingjie Wang 0002, Hui Huang 0004, Minglun Gong |
Vis. Comput. | 4 |
| 2022 | A Balanced-Partitioning Treemapping Method for Digital Hierarchical DatasetabstractThe problem of visualizing a hierarchical dataset is an important and useful technical in many real happened situations. Folder system, stock market, and other hierarchical related dataset can use this technical for better understanding the structure, dynamic variation of the dataset. Traditional space-filling(square) based methods have advantages of compact space usage, node size showing compared to diagram based methods. While space-filling based methods have two main research directions—static and dynamic performance. We present a treemapping method based on balanced partitioning that enables in one variant very good aspect ratios, in another good temporal coherence for dynamic data and in the third a good compromise between these two aspects. To layout a treemap, we divide all children of a node into two groups. These groups are further divided until we reach groups of single elements. Then these groups are combined to form the rectangle representing the parent node. This process is performed for each layer of a given hierarchical dataset. In one variant of our partitioning we sort child elements first and built two as equal as possible sized groups from big and small elements(size-balanced partition), which achieves good aspect ratios for the rectangles, but less good temporal coherence(dynamic). The second variant takes the sequence of children and creates the as equal as possible groups with-out sorting(sequence-based, good compromise between aspect ratio and temporal coherency). The third variant splits the children sets always into two groups of equal cardinality regardless of their size(number-balanced, worse aspect ratios but good temporal coherence). We evaluate aspect ratios and dynamic stability of our methods and propose a new metric that measures the visual difference between rectangles during their movement for representing temporally changing inputs. We demonstrate that our treemapping via balanced partitioning out performs state-of-the-art methods for a number of real-world datasets. Minglun Gong, Oliver Deussen |
Virtual Real. Intell. Hardw. | 2 |
| 2021 | EventHPE: Event-based 3D Human Pose and Shape EstimationabstractEvent camera is an emerging imaging sensor for capturing dynamics of moving objects as events, which motivates our work in estimating 3D human pose and shape from the event signals. Events, on the other hand, have their unique challenges: rather than capturing static body postures, the event signals are best at capturing local motions. This leads us to propose a two-stage deep learning approach, called EventHPE. The first-stage, FlowNet, is trained by unsupervised learning to infer optical flow from events. Both events and optical flow are closely related to human body dynamics, which are fed as input to the ShapeNet in the second stage, to estimate 3D human shapes. To mitigate the discrepancy between image-based flow (optical flow) and shape-based flow (vertices movement of human body shape), a novel flow coherence loss is introduced by exploiting the fact that both flows are originated from the identical human motion. An in-house event-based 3D human dataset is curated that comes with 3D pose and shape annotations, which is by far the largest one to our knowledge. Empirical evaluations on DHP19 dataset and our in-house dataset demonstrate the effectiveness of our approach. Shihao Zou, Chuan Guo 0002, Xinxin Zuo, Sen Wang 0003, Pengyu Wang 0007, Xiaoqin Hu, Shoushun Chen, Minglun Gong, Li Cheng 0001 |
ICCV | 8 |
| 2021 | Interlayer and intralayer scale aggregation for scale-invariant crowd counting
Mingjie Wang 0002, Hao Cai 0004, Jun Zhou 0023, Minglun Gong |
Neurocomputing | 4 |
| 2021 | SparseFusion: Dynamic Human Avatar Modeling From Sparse RGBD ImagesabstractIn this paper, we propose a novel approach to reconstruct 3D human body shapes based on a sparse set of RGBD frames using a single RGBD camera. We specifically focus on the realistic settings where human subjects move freely during the capture. The main challenge is how to robustly fuse these sparse frames into a canonical 3D model, under pose changes and surface occlusions. This is addressed by our new framework consisting of the following steps. First, based on a generative human template, for every two frames having sufficient overlap, an initial pairwise alignment is performed; It is followed by a global non-rigid registration procedure, in which partial results from RGBD frames are collected into a unified 3D shape, under the guidance of correspondences from the pairwise alignment; Finally, the texture map of the reconstructed human model is optimized to deliver a clear and spatially consistent texture. Empirical evaluations on synthetic and real datasets demonstrate both quantitatively and qualitatively the superior performance of our framework in reconstructing complete 3D human models with high fidelity. It is worth noting that our framework is flexible, with potential applications going beyond shape reconstruction. As an example, we showcase its use in reshaping and reposing to a new avatar. Xinxin Zuo, Sen Wang 0003, Jiangbin Zheng 0001, Minglun Gong, Ruigang Yang, Li Cheng 0001 |
IEEE Trans. Multim. | 5 |
| 2021 | Aerial path planning for online real-time exploration and offline high-quality reconstruction of large-scale urban scenesabstractExisting approaches have shown that, through carefully planning flight trajectories, images captured by Unmanned Aerial Vehicles (UAVs) can be used to reconstruct high-quality 3D models for real environments. These approaches greatly simplify and cut the cost of large-scale urban scene reconstruction. However, to properly capture height discontinuities in urban scenes, all state-of-the-art methods require prior knowledge on scene geometry and hence, additional prepossessing steps are needed before performing the actual image acquisition flights. To address this limitation and to make urban modeling techniques even more accessible, we present a real-time explore-and-reconstruct planning algorithm that does not require any prior knowledge for the scenes. Using only captured 2D images, we estimate 3D bounding boxes for buildings on-the-fly and use them to guide online path planning for both scene exploration and building observation. Experimental results demonstrate that the aerial paths planned by our algorithm in realtime for unknown environments support reconstructing 3D models with comparable qualities and lead to shorter flight air time. Ruiqi Cui, Ke Xie 0001, Minglun Gong, Hui Huang 0004 |
ACM Trans. Graph. | 4 |
| 2021 | Fine-grained talking face generation with video reinterpretation
Xin Huang 0030, Mingjie Wang 0002, Minglun Gong |
Vis. Comput. | 3 |
| 2020 | EFANet: Exchangeable Feature Alignment Network for Arbitrary Style TransferabstractStyle transfer has been an important topic both in computer vision and graphics. Since the seminal work of Gatys et al. first demonstrates the power of stylization through optimization in the deep feature space, quite a few approaches have achieved real-time arbitrary style transfer with straightforward statistic matching techniques. In this work, our key observation is that only considering features in the input style image for the global deep feature statistic matching or local patch swap may not always ensure a satisfactory style transfer; see e.g., Figure 1. Instead, we propose a novel transfer framework, EFANet, that aims to jointly analyze and better align exchangeable features extracted from the content and style image pair. In this way, the style feature from the style image seeks for the best compatibility with the content information in the content image, leading to more structured stylization results. In addition, a new whitening loss is developed for purifying the computed content features and better fusion with styles in feature space. Qualitative and quantitative experiments demonstrate the advantages of our approach. Chunjin Song, Yang Zhou 0007, Minglun Gong, Hui Huang 0004 |
AAAI | 4 |
| 2020 | 3D Human Shape Reconstruction from a Polarization Image
Shihao Zou, Xinxin Zuo, Yiming Qian, Sen Wang 0003, Chi Xu 0002, Minglun Gong, Li Cheng 0001 |
ECCV (14) | 6 |
| 2020 | Stochastic Multi-Scale Aggregation Network for Crowd CountingabstractCrowd counting from unconstrained and congested scenes is an important task in computer vision. Its main difficulties stem from large scale/density variation and prone to over-fitting. This paper presents a novel end-to-end stochastic multi-scale aggregation network (SMANet) which carefully addresses these issues. Specifically, general features are first extracted by the front-end subnetwork and then fed into the back-end subnetwork which consists of stochastic multi-scale aggregation module, density map generator, and global prior encoder. The stochastic aggregation impels the multi-branch units to learn features at different scales effectively and reduces sensitivity to scale variations, whereas the global prior encoder is designed to encode global contextual information and guarantee density consistency of shared representations. Our proposed SMANet is the first work to fuse multi-scale features in a stochastic manner for crowd counting. Experimental results on four public datasets demonstrate that our SMANet consistently outperforms the state-of-the-arts. Mingjie Wang 0002, Hao Cai 0004, Jun Zhou 0023, Minglun Gong |
ICASSP | 4 |
| 2020 | Action2Motion: Conditioned Generation of 3D Human MotionsabstractAction recognition is a relatively established task, where given an input sequence of human motion, the goal is to predict its action category. This paper, on the other hand, considers a relatively new problem, which could be thought of as an inverse of action recognition: given a prescribed action type, we aim to generate plausible human motion sequences in 3D. Importantly, the set of generated motions are expected to maintain its diversity to be able to explore the entire action-conditioned motion space; meanwhile, each sampled sequence faithfully resembles a natural human body articulation dynamics. Motivated by these objectives, we follow the physics law of human kinematics by adopting the Lie Algebra theory to represent the natural human motions; we also propose a temporal Variational Auto-Encoder (VAE) that encourages a diverse sampling of the motion space. A new 3D human motion dataset, HumanAct12, is also constructed. Empirical experiments over three distinct human motion datasets (including ours) demonstrate the effectiveness of our approach. Chuan Guo 0002, Xinxin Zuo, Sen Wang 0003, Shihao Zou, Qingyao Sun, Annan Deng, Minglun Gong, Li Cheng 0001 |
ACM Multimedia | 7 |
| 2020 | ADNet: Adaptively Dense Convolutional Neural NetworksabstractConvolutional neural networks (CNNs) have demonstrated great success in vision tasks. However, most existing architectures still suffer from low feature reuse efficiency. In this paper, we present a layer attention based Adaptively Dense Network (ADNet) by adaptively determining the reuse status of hierarchical preceding features. Specifically, a dense residual aggregation strategy is developed to fuse multi-level internal representations in an effective manner. Furthermore, a novel layer attention mechanism is proposed to explicitly model the interrelationship among layers to automatically adjust the density of the network. It is worth noting that existing ResNets and DenseNets are both special cases of our ADNet. Extensive experiments demonstrate that the proposed architecture consistently and indubitably achieves competitive results in accuracy on benchmark datasets (CIFAR10, CIFAR100, and SVHN), while at the same time remarkably reduces computational costs and memory space. Visualization and analysis on layer-wise attention further provide better understanding on the density of feature reuse in Deep Networks. Mingjie Wang 0002, Hao Cai 0004, Xin Huang 0030, Minglun Gong |
WACV | 4 |
| 2020 | SiamesePointNet: A Siamese Point Network Architecture for Learning 3D Shape DescriptorabstractAbstract We present a novel deep learning approach to extract point‐wise descriptors directly on 3D shapes by introducing Siamese Point Networks, which contain a global shape constraint module and a feature transformation operator. Such geometric descriptor can be used in a variety of shape analysis problems such as 3D shape dense correspondence, key point matching and shape‐to‐scan matching. The descriptor is produced by a hierarchical encoder–decoder architecture that is trained to map geometrically and semantically similar points close to one another in descriptor space. Benefiting from the additional shape contrastive constraint and the hierarchical local operator, the learned descriptor is highly aware of both the global context and local context. In addition, a feature transformation operation is introduced in the end of our networks to transform the point features to a compact descriptor space. The feature transformation can make the descriptors extracted by our networks unaffected by geometric differences in shapes. Finally, an N‐tuple loss is used to train all the point descriptors on a complete 3D shape simultaneously to obtain point‐wise descriptors. The proposed Siamese Point Networks are robust to many types of perturbations such as the Gaussian noise and partial scan. In addition, we demonstrate that our approach improves state‐of‐the‐art results on the BHCP benchmark. M. J. Wang, W. D. Mao, Minglun Gong, X. P. Liu |
Comput. Graph. Forum | 4 |
| 2020 | No-reference image sharpness assessment based on discrepancy measures of structural degradation
Hao Cai 0004, Mingjie Wang 0002, Wendong Mao, Minglun Gong |
J. Vis. Commun. Image Represent. | 4 |
| 2020 | BSD-GAN: Branched Generative Adversarial Network for Scale-Disentangled Representation Learning and Image SynthesisabstractWe introduce BSD-GAN, a novel multi-branch and scale-disentangled training method which enables unconditional Generative Adversarial Networks (GANs) to learn image representations at multiple scales, benefiting a wide range of generation and editing tasks. The key feature of BSD-GAN is that it is trained in multiple branches, progressively covering both the breadth and depth of the network, as resolutions of the training images increase to reveal finer-scale features. Specifically, each noise vector, as input to the generator network of BSD-GAN, is deliberately split into several sub-vectors, each corresponding to, and is trained to learn, image representations at a particular scale. During training, we progressively "de-freeze" the sub-vectors, one at a time, as a new set of higher-resolution images is employed for training and more network layers are added. A consequence of such an explicit sub-vector designation is that we can directly manipulate and even combine latent (sub-vector) codes which model different feature scales. Extensive experiments demonstrate the effectiveness of our training method in scale-disentangled learning of image representations and synthesis of novel image contents, without any extra labels and without compromising quality of the synthesized high-resolution images. We further demonstrate several image generation and manipulation applications enabled or improved by BSD-GAN. Zili Yi, Hao Cai 0004, Wendong Mao, Minglun Gong, Hao (Richard) Zhang |
IEEE Trans. Image Process. | 5 |
| 2020 | TAP-Net: transport-and-pack using reinforcement learningabstractWe introduce the transport-and-pack (TAP) problem, a frequently encountered instance of real-world packing, and develop a neural optimization solution based on reinforcement learning. Given an initial spatial configuration of boxes, we seek an efficient method to iteratively transport and pack the boxes compactly into a target container. Due to obstruction and accessibility constraints, our problem has to add a new search dimension, i.e., finding an optimal transport sequence , to the already immense search space for packing alone. Using a learning-based approach, a trained network can learn and encode solution patterns to guide the solution of new problem instances instead of executing an expensive online search. In our work, we represent the transport constraints using a precedence graph and train a neural network, coined TAP-Net, using reinforcement learning to reward efficient and stable packing. The network is built on an encoder-decoder architecture, where the encoder employs convolution layers to encode the box geometry and precedence graph and the decoder is a recurrent neural network (RNN) which inputs the current encoder output, as well as the current box packing state of the target container, and outputs the next box to pack, as well as its orientation. We train our network on randomly generated initial box configurations, without supervision , via policy gradients to learn optimal TAP policies to maximize packing efficiency and stability. We demonstrate the performance of TAP-Net on a variety of examples, evaluating the network through ablation studies and comparisons to baselines and alternative network designs. We also show that our network generalizes well to larger problem instances, when trained on small-sized inputs. Ruizhen Hu, Juzhan Xu, Minglun Gong, Hao (Richard) Zhang, Hui Huang 0004 |
ACM Trans. Graph. | 4 |
| 2020 | Offsite aerial path planning for efficient urban scene reconstructionabstractWith rapid development in UAV technologies, it is now possible to reconstruct large-scale outdoor scenes using only images captured by low-cost drones. The problem, however, becomes how to plan the aerial path for a drone to capture images so that two conflicting goals are optimized: maximizing the reconstruction quality and minimizing mid-air image acquisition effort. Existing approaches either resort to pre-defined dense and thus inefficient view sampling strategy, or plan the path adaptively but require two onsite flight passes and intensive computation in-between. Hence, using these methods to capture and reconstruct large-scale scenes can be tedious. In this paper, we present an adaptive aerial path planning algorithm that can be done before the site visit. Using only a 2D map and a satellite image of the to-be-reconstructed area, we first compute a coarse 2.5D model for the scene based on the relationship between buildings and their shadows. A novel Max-Min optimization is then proposed to select a minimal set of viewpoints that maximizes the reconstructability under the the same number of viewpoints. Experimental results on benchmark show that our planning approach can effectively reduce the number of viewpoints needed than the previous state-of-the-art method, while maintaining comparable reconstruction quality. Since no field computation or a second visit is needed, and the view number is also minimized, our approach significantly reduces the time required in the field as well as the off-line computation cost for multi-view stereo reconstruction, making it possible to reconstruct a large-scale urban scene in a short time with moderate effort. Ke Xie 0001, Yang Zhou 0007, Minglun Gong, Hui Huang 0004 |
ACM Trans. Graph. | 6 |
| 2019 | A Global-Matching Framework for Multi-View Stereopsis
Wendong Mao, Minglun Gong, Xin Huang 0030, Hao Cai 0004, Zili Yi |
CAIP (1) | 2 |
| 2019 | ETNet: Error Transition Network for Arbitrary Style TransferabstractNumerous valuable efforts have been devoted to achieving arbitrary style transfer since the seminal work of Gatys et al. However, existing state-of-the-art approaches often generate insufficiently stylized results under challenging cases. We believe a fundamental reason is that these approaches try to generate the stylized result in a single shot and hence fail to fully satisfy the constraints on semantic structures in the content images and style patterns in the style images. Inspired by the works on error-correction, instead, we propose a self-correcting model to predict what is wrong with the current stylization and refine it accordingly in an iterative manner. For each refinement, we transit the error features across both the spatial and scale domain and invert the processed features into a residual image, with a network we call Error Transition Network (ETNet). The proposed model improves over the state-of-the-art methods with better semantic structures and more adaptive style pattern details. Various qualitative and quantitative experiments show that the key concept of both progressive strategy and error-correction leads to better results. Code and models are available at https://github.com/zhijieW94/ETNet. Chunjin Song, Yang Zhou 0007, Minglun Gong, Hui Huang 0004 |
NeurIPS | 4 |
| 2019 | Semi-Dense Stereo Matching Using Dual CNNsabstractA robust solution for semi-dense stereo matching is presented. It utilizes two CNN models for computing stereo matching cost and performing confidence-based filtering, respectively. Compared to existing CNNs-based matching cost generation approaches, our method feeds additional global information into the network so that the learned model can better handle challenging cases, such as lighting changes and lack of textures. Through utilizing non-parametric transforms, our method is also more self-reliant than most existing semi-dense stereo approaches, which rely highly on the adjustment of parameters. The experimental results based on Middlebury Stereo dataset demonstrate that the proposed approach outperforms the state-of-the-art semi-dense stereo approaches. Wendong Mao, Mingjie Wang 0002, Jun Zhou 0023, Minglun Gong |
WACV | 4 |
| 2019 | Multi-Scale Convolution Aggregation and Stochastic Feature Reuse for DenseNetsabstractRecently, Convolution Neural Networks (CNNs) obtained huge success in numerous vision tasks. In particular, DenseNets have demonstrated that feature reuse via dense skip connections can effectively alleviate the difficulty of training very deep networks and that reusing features generated by the initial layers in all subsequent layers has strong impact on performance. To feed even richer information into the network, a novel adaptive Multi-scale Convolution Aggregation module is presented in this paper. Composed of layers for multi-scale convolutions, trainable cross-scale aggregation, maxout, and concatenation, this module is highly non-linear and can boost the accuracy of DenseNet while using much fewer parameters. In addition, due to high model complexity, the network with extremely dense feature reuse is prone to overfitting. To address this problem, a regularization method named Stochastic Feature Reuse is also presented. Through randomly dropping a set of feature maps to be reused for each mini-batch during the training phase, this regularization method reduces training costs and prevents co-adaptation. Experimental results on CIFAR-10, CIFAR-100 and SVHN benchmarks demonstrated the effectiveness of the proposed methods. Mingjie Wang 0002, Jun Zhou 0023, Wendong Mao, Minglun Gong |
WACV | 4 |
| 2019 | Blind quality assessment of gamut-mapped images via local and global statistical analysis
Hao Cai 0004, Leida Li, Zili Yi, Minglun Gong |
J. Vis. Commun. Image Represent. | 4 |
| 2019 | Towards a blind image quality evaluator using multi-scale second-order statistics
Hao Cai 0004, Leida Li, Zili Yi, Minglun Gong |
Signal Process. Image Commun. | 4 |
| 2019 | A unified framework for exploring time-varying volumetric data based on block correspondenceabstractEffective exploration of spatiotemporal volumetric data sets remains a key challenge in scientific visualization. Although great advances have been made over the years, existing solutions typically focus on only one or two aspects of data analysis and visualization. A streamlined workflow for analyzing time-varying data in a comprehensive and unified manner is still missing. Towards this goal, we present a novel approach for time-varying data visualization that encompasses keyframe identification, feature extraction and tracking under a single, unified framework. At the heart of our approach lies in the GPU-accelerated BlockMatch method, a dense block correspondence technique that extends the PatchMatch method from 2D pixels to 3D voxels. Based on the results of dense correspondence, we are able to identify keyframes from the time sequence using k-medoids clustering along with a bidirectional similarity measure. Furthermore, in conjunction with the graph cut algorithm, this framework enables us to perform fine-grained feature extraction and tracking. We tested our approach using several time-varying data sets to demonstrate its effectiveness and utility. Kecheng Lu 0002, Chaoli Wang 0001, Keqin Wu, Minglun Gong, Yunhai Wang |
Vis. Informatics | 4 |
| 2018 | Simultaneous 3D Reconstruction for Water Surface and Underwater Scene
Yiming Qian, Yinqiang Zheng, Minglun Gong, Yee-Hong Yang |
ECCV (3) | 3 |
| 2018 | Structure-Aware Data ConsolidationabstractWe present a structure-aware technique to consolidate noisy data, which we use as a pre-process for standard clustering and dimensionality reduction. Our technique is related to mean shift, but instead of seeking density modes, it reveals and consolidates continuous high density structures such as curves and surface sheets in the underlying data while ignoring noise and outliers. We provide a theoretical analysis under a Gaussian noise model, and show that our approach significantly improves the performance of many non-linear dimensionality reduction and clustering algorithms in challenging scenarios. Peter Bertholet, Hui Huang 0004, Daniel Cohen-Or, Minglun Gong, Matthias Zwicker |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2018 | Multi-modal feature fusion for geographic image annotation
Ke Li 0005, Changqing Zou, Shuhui Bu, Yun Liang 0003, Jian Zhang 0026, Minglun Gong |
Pattern Recognit. | 6 |
| 2018 | Appearance Modeling via Proxy-to-Image AlignmentabstractEndowing 3D objects with realistic surface appearance is a challenging and time-demanding task, as real-world surfaces typically exhibit a plethora of spatially variant geometric and photometric detail. Not surprisingly, computer artists commonly use images of real-world objects as an inspiration and a reference for their digital creations. However, despite two decades of research on image-based modeling, there are still no tools available for automatically extracting the detailed appearance (microgeometry and texture) of a 3D surface from a single image. In this article, we present a novel user-assisted approach for quickly and easily extracting a nonparametric appearance model from a single photograph of a reference object. The extraction process requires a user-provided proxy, whose geometry roughly approximates that of the object in the image. Since the proxy is just a rough approximation, it is necessary to align and deform it so as to match the reference object. The main contribution of this work is a novel technique to perform such an alignment, which enables accurate joint recovery of geometric detail and reflectance. The correlations between the recovered geometry at various scales and the spatially varying reflectance constitute a nonparametric appearance model. Once extracted, the appearance model may then be applied to various 3D shapes, whose large-scale geometry may differ considerably from that of the original reference object. Thus, our approach makes it possible to construct an appearance library, allowing users to easily enrich detail-less 3D shapes with realistic geometric detail and surface texture. Hui Huang 0004, Ke Xie 0001, Dani Lischinski, Minglun Gong, Xin Tong 0001, Daniel Cohen-Or |
ACM Trans. Graph. | 5 |
| 2018 | Full 3D reconstruction of transparent objectsabstractNumerous techniques have been proposed for reconstructing 3D models for opaque objects in past decades. However, none of them can be directly applied to transparent objects. This paper presents a fully automatic approach for reconstructing complete 3D shapes of transparent objects. Through positioning an object on a turntable, its silhouettes and light refraction paths under different viewing directions are captured. Then, starting from an initial rough model generated from space carving, our algorithm progressively optimizes the model under three constraints: surface and refraction normal consistency, surface projection and silhouette consistency, and surface smoothness. Experimental results on both synthetic and real objects demonstrate that our method can successfully recover the complex shapes of transparent objects and faithfully reproduce their light refraction properties. Bojian Wu, Yang Zhou 0007, Yiming Qian, Minglun Gong, Hui Huang 0004 |
ACM Trans. Graph. | 4 |
| 2018 | Creating and chaining camera moves for quadrotor videographyabstractCapturing aerial videos with a quadrotor-mounted camera is a challenging creative task, as it requires the simultaneous control of the quadrotor's motion and the mounted camera's orientation. Letting the drone follow a pre-planned trajectory is a much more appealing option, and recent research has proposed a number of tools designed to automate the generation of feasible camera motion plans; however, these tools typically require the user to specify and edit the camera path, for example by providing a complete and ordered sequence of key viewpoints. In this paper, we propose a higher level tool designed to enable even novice users to easily capture compelling aerial videos of large-scale outdoor scenes. Using a coarse 2.5D model of a scene, the user is only expected to specify starting and ending viewpoints and designate a set of landmarks, with or without a particular order. Our system automatically generates a diverse set of candidate local camera moves for observing each landmark, which are collision-free, smooth, and adapted to the shape of the landmark. These moves are guided by a landmark-centric view quality field, which combines visual interest and frame composition. An optimal global camera trajectory is then constructed that chains together a sequence of local camera moves, by choosing one move for each landmark and connecting them with suitable transition trajectories. This task is formulated and solved as an instance of the Set Traveling Salesman Problem. Ke Xie 0001, Shengqiu Huang, Dani Lischinski, Marc Christie, Kai Xu 0004, Minglun Gong, Daniel Cohen-Or, Hui Huang 0004 |
ACM Trans. Graph. | 7 |
| 2017 | Stereo-Based 3D Reconstruction of Dynamic Fluid Surfaces by Global Optimizationabstract3D Reconstruction of dynamic fluid surfaces is an open and challenging problem in computer vision. Unlike previous approaches that reconstruct each surface point independently and often return noisy depth maps, we propose a novel global optimization-based approach that recovers both depths and normals of all 3D points simultaneously. Using the traditional refraction stereo setup, we capture the wavy appearance of a pre-generated random pattern, and then estimate the correspondences between the captured images and the known background by tracking the pattern. Assuming that the light is refracted only once through the fluid interface, we minimize an objective function that incorporates both the cross-view normal consistency constraint and the single-view normal consistency constraints. The key idea is that the normals required for light refraction based on Snells law from one view should agree with not only the ones from the second view, but also the ones estimated from local 3D geometry. Moreover, an effective reconstruction error metric is designed for estimating the refractive index of the fluid. We report experimental results on both synthetic and real data demonstrating that the proposed approach is accurate and shows superiority over the conventional stereo-based method. Yiming Qian, Minglun Gong, Yee-Hong Yang |
CVPR | 2 |
| 2017 | DualGAN: Unsupervised Dual Learning for Image-to-Image TranslationabstractConditional Generative Adversarial Networks (GANs) for cross-domain image-to-image translation have made much progress recently [7, 8, 21, 12, 4, 18]. Depending on the task complexity, thousands to millions of labeled image pairs are needed to train a conditional GAN. However, human labeling is expensive, even impractical, and large quantities of data may not always be available. Inspired by dual learning from natural language translation [23], we develop a novel dual-GAN mechanism, which enables image translators to be trained from two sets of unlabeled images from two domains. In our architecture, the primal GAN learns to translate images from domain U to those in domain V, while the dual GAN learns to invert the task. The closed loop made by the primal and dual tasks allows images from either domain to be translated and then reconstructed. Hence a loss function that accounts for the reconstruction error of images can be used to train the translators. Experiments on multiple image translation tasks with unlabeled data show considerable performance gain of DualGAN over a single GAN. For some tasks, DualGAN can even achieve comparable or slightly better results than conditional GAN trained on fully labeled data. Zili Yi, Hao (Richard) Zhang, Ping Tan 0002, Minglun Gong |
ICCV | 4 |
| 2017 | Linear Discriminative Star Coordinates for Exploring Class and Cluster Separation of High Dimensional DataabstractAbstract One main task for domain experts in analysing their nD data is to detect and interpret class/cluster separations and outliers. In fact, an important question is, which features/dimensions separate classes best or allow a cluster‐based data classification. Common approaches rely on projections from nD to 2D, which comes with some challenges, such as: The space of projection contains an infinite number of items. How to find the right one? The projection approaches suffers from distortions and misleading effects. How to rely to the projected class/cluster separation? The projections involve the complete set of dimensions/features. How to identify irrelevant dimensions? Thus, to address these challenges, we introduce a visual analytics concept for the feature selection based on linear discriminative star coordinates (DSC), which generate optimal cluster separating views in a linear sense for both labeled and unlabeled data. This way the user is able to explore how each dimension contributes to clustering. To support to explore relations between clusters and data dimensions, we provide a set of cluster‐aware interactions allowing to smartly iterate through subspaces of both records and features in a guided manner. We demonstrate our features selection approach for optimal cluster/class separation analysis with a couple of experiments on real‐life benchmark high‐dimensional data sets. Yunhai Wang, Feiping Nie 0001, Holger Theisel, Minglun Gong, Dirk J. Lehmann |
Comput. Graph. Forum | 5 |
| 2017 | 4D Reconstruction of Blooming FlowersabstractAbstract Flower blooming is a beautiful phenomenon in nature as flowers open in an intricate and complex manner whereas petals bend, stretch and twist under various deformations. Flower petals are typically thin structures arranged in tight configurations with heavy self‐occlusions. Thus, capturing and reconstructing spatially and temporally coherent sequences of blooming flowers is highly challenging. Early in the process only exterior petals are visible and thus interior parts will be completely missing in the captured data. Utilizing commercially available 3D scanners, we capture the visible parts of blooming flowers into a sequence of 3D point clouds. We reconstruct the flower geometry and deformation over time using a template‐based dynamic tracking algorithm. To track and model interior petals hidden in early stages of the blooming process, we employ an adaptively constrained optimization. Flower characteristics are exploited to track petals both forward and backward in time. Our methods allow us to faithfully reconstruct the flower blooming process of different species. In addition, we provide comparisons with state‐of‐the‐art physical simulation‐based approaches and evaluate our approach by using photos of captured real flowers. Xiaochen Fan, Minglun Gong, Andrei Sharf, Oliver Deussen, Hui Huang 0004 |
Comput. Graph. Forum | 3 |
| 2017 | Analysis and Controlled Synthesis of Inhomogeneous TexturesabstractMany interesting real-world textures are inhomogeneous and/or anisotropic. An inhomogeneous texture is one where various visual properties exhibit significant changes across the texture's spatial domain. Examples include perceptible changes in surface color, lighting, local texture pattern and/or its apparent scale, and weathering effects, which may vary abruptly, or in a continuous fashion. An anisotropic texture is one where the local patterns exhibit a preferred orientation, which also may vary across the spatial domain. While many example-based texture synthesis methods can be highly effective when synthesizing uniform (stationary) isotropic textures, synthesizing highly non-uniform textures, or ones with spatially varying orientation, is a considerably more challenging task, which so far has remained underexplored. In this paper, we propose a new method for automatic analysis and controlled synthesis of such textures. Given an input texture exemplar, our method generates a source guidance map comprising: (i) a scalar progression channel that attempts to capture the low frequency spatial changes in color, lighting, and local pattern combined, and (ii) a direction field that captures the local dominant orientation of the texture. Having augmented the texture exemplar with this guidance map, users can exercise better control over the synthesized result by providing easily specified target guidance maps, which are used to constrain the synthesis process. Yang Zhou 0007, Huajie Shi, Dani Lischinski, Minglun Gong, Johannes Kopf 0001, Hui Huang 0004 |
Comput. Graph. Forum | 4 |
| 2017 | Unsupervised hierarchical image segmentation through fuzzy entropy maximization
Shibai Yin, Yiming Qian, Minglun Gong |
Pattern Recognit. | 3 |
| 2017 | Artistic stylization of face photos based on a single exemplar
Zili Yi, Songyuan Ji, Minglun Gong |
Vis. Comput. | 4 |
| 2016 | 3D Reconstruction of Transparent Objects with Position-Normal ConsistencyabstractEstimating the shape of transparent and refractive objects is one of the few open problems in 3D reconstruction. Under the assumption that the rays refract only twice when traveling through the object, we present the first approach to simultaneously reconstructing the 3D positions and normals of the object's surface at both refraction locations. Our acquisition setup requires only two cameras and one monitor, which serves as the light source. After acquiring the ray-ray correspondences between each camera and the monitor, we solve an optimization function which enforces a new position-normal consistency constraint. That is, the 3D positions of surface points shall agree with the normals required to refract the rays under Snell's law. Experimental results using both synthetic and real data demonstrate the robustness and accuracy of the proposed approach. Yiming Qian, Minglun Gong, Yee-Hong Yang |
CVPR | 2 |
| 2016 | Artificial Multi-Bee-Colony Algorithm for k-Nearest-Neighbor Fields SearchabstractSearching the k-nearest matching patches for each patch in an input image, i.e., computing the k-nearest-neighbor fields ($k$-NNF), is a core part of various computer vision/graphics algorithms. In this paper, we show that $k$-NNF can be efficiently computed using a novel artificial multi-bee-colony (AMBC) algorithm, where each patch uses a dedicated bee colony to search for its k-nearest matches. As a population-based algorithm, AMBC is capable of escaping local optima. The added communication among different colonies further allows good matches to be quickly propagated across the image. In addition, AMBC makes no assumption about the neighborhood structure or communication direction, making it directly applicable to image sets and suitable for parallel processing. Quantitative evaluations show that AMBC can find solutions that are much closer to the ground truth than the generalized PatchMatch algorithm does. It also outperforms the PatchMatch Graph over image sets. Yunhai Wang, Yiming Qian, Minglun Gong, Wolfgang Banzhaf |
GECCO | 4 |
| 2016 | Trip Synopsis: 60km in 60secabstractAbstract Computerized route planning tools are widely used today by travelers all around the globe, while 3D terrain and urban models are becoming increasingly elaborate and abundant. This makes it feasible to generate a virtual 3D flyby along a planned route. Such a flyby may be useful, either as a preview of the trip, or as an after‐the‐fact visual summary. However, a naively generated preview is likely to contain many boring portions, while skipping too quickly over areas worthy of attention. In this paper, we introduce 3D trip synopsis: a continuous visual summary of a trip that attempts to maximize the total amount of visual interest seen by the camera. The main challenge is to generate a synopsis of a prescribed short duration, while ensuring a visually smooth camera motion. Using an application‐specific visual interest metric, we measure the visual interest at a set of viewpoints along an initial camera path, and maximize the amount of visual interest seen in the synopsis by varying the speed along the route. A new camera path is then computed using optimization to simultaneously satisfy requirements, such as smoothness, focus and distance to the route. The process is repeated until convergence. The main technical contribution of this work is a new camera control method, which iteratively adjusts the camera trajectory and determines all of the camera trajectory parameters, including the camera position, altitude, heading, and tilt. Our results demonstrate the effectiveness of our trip synopses, compared to a number of alternatives. Hui Huang 0004, Dani Lischinski, Hao (Richard) Zhang, Minglun Gong, Marc Christie, Daniel Cohen-Or |
Comput. Graph. Forum | 4 |
| 2016 | Full 3D Plant Reconstruction via Intrusive AcquisitionabstractAbstract Digitally capturing vegetation using off‐the‐shelf scanners is a challenging problem. Plants typically exhibit large self‐occlusions and thin structures which cannot be properly scanned. Furthermore, plants are essentially dynamic, deforming over the time, which yield additional difficulties in the scanning process. In this paper, we present a novel technique for acquiring and modelling of plants and foliage. At the core of our method is an intrusive acquisition approach, which disassembles the plant into disjoint parts that can be accurately scanned and reconstructed offline. We use the reconstructed part meshes as 3D proxies for the reconstruction of the complete plant and devise a global‐to‐local non‐rigid registration technique that preserves specific plant characteristics. Our method is tested on plants of various styles, appearances and characteristics. Results show successful reconstructions with high accuracy with respect to the acquired data. Kangxue Yin, Hui Huang 0004, Pinxin Long, Alexei Gaissinski, Minglun Gong, Andrei Sharf |
Comput. Graph. Forum | 5 |
| 2015 | Frequency-Based Environment Matting by Compressive SensingabstractExtracting environment mattes using existing approaches often requires either thousands of captured images or a long processing time, or both. In this paper, we propose a novel approach to capturing and extracting the matte of a real scene effectively and efficiently. Grown out of the traditional frequency-based signal analysis, our approach can accurately locate contributing sources. By exploiting the recently developed compressive sensing theory, we simplify the data acquisition process of frequency-based environment matting. Incorporating phase information in a frequency signal into data acquisition further accelerates the matte extraction procedure. Compared with the state-of-the-art method, our approach achieves superior performance on both synthetic and real data, while consuming only a fraction of the processing time. Yiming Qian, Minglun Gong, Yee-Hong Yang |
ICCV | 2 |
| 2015 | Recovering intrinsic images from image sequences using total variation modelsabstractRecovering intrinsic images from natural photos is one of the foundational problems in computer vision. This mission always falls into an ill-posed problem. In order to attain reasonable estimations, one strategy is to use multiple images of the scene under various lightings so as to narrow the solution space, whereas another is to utilize priori knowledge as constraints. In this paper, we present an approach to deriving intrinsic images (including illumination images and reflectance images) that employs both strategies. Specifically, the Total Variation (TV) constraint is imposed because of its excellent edge preservation ability and simple parameter settings. To solve this objective function efficiently, we propose using the Alternating Direction Method of Multipliers (AD-MM) to build an iterative numerical scheme. Experimental results illustrate the effectiveness of the proposed model and the numerical scheme. Xiaohua Xie, Wenyong Gong, Minglun Gong, Tieru Wu |
ICIP | 3 |
| 2015 | Distilled Collections from Textual Image QueriesabstractAbstract We present a distillation algorithm which operates on a large, unstructured, and noisy collection of internet images returned from an online object query. We introduce the notion of a distilled set, which is a clean, coherent, and structured subset of inlier images. In addition, the object of interest is properly segmented out throughout the distilled set. Our approach is unsupervised, built on a novel clustering scheme, and solves the distillation and object segmentation problems simultaneously. In essence, instead of distilling the collection of images, we distill a collection of loosely cutout foreground “shapes”, which may or may not contain the queried object. Our key observation, which motivated our clustering scheme, is that outlier shapes are expected to be random in nature, whereas, inlier shapes, which do tightly enclose the object of interest, tend to be well supported by similar shapes captured in similar views. We analyze the commonalities among candidate foreground segments, without aiming to analyze their semantics, but simply by clustering similar shapes and considering only the most significant clusters representing non‐trivial shapes. We show that when tuned conservatively, our distillation algorithm is able to extract a near perfect subset of true inliers. Furthermore, we show that our technique scales well in the sense that the precision rate remains high, as the collection grows. We demonstrate the utility of our distillation results with a number of interesting graphics applications. Hadar Averbuch-Elor, Yunhai Wang, Yiming Qian, Minglun Gong, Johannes Kopf 0001, Hao (Richard) Zhang, Daniel Cohen-Or |
Comput. Graph. Forum | 4 |
| 2015 | Integrated Foreground Segmentation and Boundary Matting for Live VideosabstractThe objective of foreground segmentation is to extract the desired foreground object from input videos. Over the years, there have been significant amount of efforts on this topic. Nevertheless, there still lacks a simple yet effective algorithm that can process live videos of objects with fuzzy boundaries (e.g., hair) captured by freely moving cameras. This paper presents an algorithm toward this goal. The key idea is to train and maintain two competing one-class support vector machines at each pixel location, which model local color distributions for both foreground and background, respectively. The usage of two competing local classifiers, as we have advocated, provides higher discriminative power while allowing better handling of ambiguities. By exploiting this proposed machine learning technique, and by addressing both foreground segmentation and boundary matting problems in an integrated manner, our algorithm is shown to be particularly competent at processing a wide range of videos with complex backgrounds from freely moving cameras. This is usually achieved with minimum user interactions. Furthermore, by introducing novel acceleration techniques and by exploiting the parallel structure of the algorithm, near real-time processing speed (14 frames/s without matting and 8 frames/s with matting on a midrange PC & GPU) is achieved for VGA-sized videos. Minglun Gong, Yiming Qian, Li Cheng 0001 |
IEEE Trans. Image Process. | 1 |
| 2015 | Deep points consolidationabstractIn this paper, we present a consolidation method that is based on a new representation of 3D point sets. The key idea is to augment each surface point into a deep point by associating it with an inner point that resides on the meso-skeleton, which consists of a mixture of skeletal curves and sheets. The deep points representation is a result of a joint optimization applied to both ends of the deep points. The optimization objective is to fairly distribute the end points across the surface and the meso-skeleton, such that the deep point orientations agree with the surface normals. The optimization converges where the inner points form a coherent meso-skeleton, and the surface points are consolidated with the missing regions completed. The strength of this new representation stems from the fact that it is comprised of both local and non-local geometric information. We demonstrate the advantages of the deep points consolidation technique by employing it to consolidate and complete noisy point-sampled geometry with large missing parts. Hui Huang 0004, Minglun Gong, Matthias Zwicker, Daniel Cohen-Or |
ACM Trans. Graph. | 3 |
| 2015 | Generalized cylinder decompositionabstractDecomposing a complex shape into geometrically simple primitives is a fundamental problem in geometry processing. We are interested in a shape decomposition problem where the simple primitives sought are generalized cylinders , which are ubiquitous in both organic forms and man-made artifacts. We introduce a quantitative measure of cylindricity for a shape part and develop a cylindricity-driven optimization algorithm, with a global objective function, for generalized cylinder decomposition. As a measure of geometric simplicity and following the minimum description length principle, cylindricity is defined as the cost of representing a cylinder through skeletal and cross-section profile curves. Our decomposition algorithm progressively builds local to non-local cylinders, which form over-complete covers of the input shape. The over-completeness of the cylinder covers ensures a conservative buildup of the cylindrical parts, leaving the final decision on decomposition to global optimization. We solve the global optimization by finding an exact cover, which optimizes the global objective function. We demonstrate results of our optimal decomposition algorithm on numerous examples and compare with other alternatives. Yang Zhou 0007, Kangxue Yin, Hui Huang 0004, Hao (Richard) Zhang, Minglun Gong, Daniel Cohen-Or |
ACM Trans. Graph. | 5 |
| 2014 | Mobility-Trees for Indoor Scenes ManipulationabstractAbstract In this work, we introduce the ‘mobility‐tree’ construct for high‐level functional representation of complex 3D indoor scenes. In recent years, digital indoor scenes are becoming increasingly popular, consisting of detailed geometry and complex functionalities. These scenes often consist of objects that reoccur in various poses and interrelate with each other. In this work we analyse the reoccurrence of objects in the scene and automatically detect their functional mobilities. ‘Mobility’ analysis denotes the motion capabilities (i.e. degree of freedom) of an object and its subpart which typically relates to their indoor functionalities. We compute an object's mobility by analysing its spatial arrangement, repetitions and relations with other objects and store it in a ‘mobility‐tree’. Repetitive motions in the scenes are grouped in ‘mobility‐groups’, for which we develop a set of sophisticated controllers facilitating semantical high‐level editing operations. We show applications of our mobility analysis to interactive scene manipulation and reorganization, and present results for a variety of indoor scenes. Andrei Sharf, Hui Huang 0004, Cheng Liang 0005, Jiapei Zhang, Baoquan Chen, Minglun Gong |
Comput. Graph. Forum | 6 |
| 2014 | Flower reconstruction from a single photoabstractAbstract We present a semi‐automatic method for reconstructing flower models from a single photograph. Such reconstruction is challenging since the 3D structure of a flower can appear ambiguous in projection. However, the flower head typically consists of petals embedded in 3D space that share similar shapes and form certain level of regular structure. Our technique employs these assumptions by first fitting a cone and subsequently a surface of revolution to the flower structure and then computing individual petal shapes from their projection in the photo. Flowers with multiple layers of petals are handled through processing different layers separately. Occlusions are dealt with both within and between petal layers. We show that our method allows users to quickly generate a variety of realistic 3D flowers from photographs and to animate an image using the underlying models reconstructed from our method. Feilong Yan, Minglun Gong, Daniel Cohen-Or, Oliver Deussen, Baoquan Chen |
Comput. Graph. Forum | 2 |
| 2014 | Efficient multilevel image segmentation through fuzzy entropy maximization and graph cut optimization
Shibai Yin, Xiangmo Zhao, Weixing Wang 0001, Minglun Gong |
Pattern Recognit. | 4 |
| 2014 | Self-Sorting Map: An Efficient Algorithm for Presenting Multimedia Data in Structured LayoutsabstractThis paper presents the Self-Sorting Map (SSM), a novel algorithm for organizing and presenting multimedia data. Given a set of data items and a dissimilarity measure between each pair of them, the SSM places each item into a unique cell of a structured layout, where the most related items are placed together and the unrelated ones are spread apart. The algorithm integrates ideas from dimension reduction, sorting, and data clustering algorithms. Instead of solving the continuous optimization problem that other dimension reduction approaches do, the SSM transforms it into a discrete labeling problem. As a result, it can organize a set of data into a structured layout without overlap, providing a simple and intuitive presentation. The algorithm is designed for sorting all data items in parallel, making it possible to arrange millions of items in seconds. Experiments on different types of data demonstrate the SSM's versatility in a variety of applications, ranging from positioning city names by proximities to presenting images according to visual similarities, to visualizing semantic relatedness between Wikipedia articles. Grant Strong, Minglun Gong |
IEEE Trans. Multim. | 2 |
| 2014 | Quality-driven poisson-guided autoscanningabstractWe present a quality-driven, Poisson-guided autonomous scanning method. Unlike previous scan planning techniques, we do not aim to minimize the number of scans needed to cover the object's surface, but rather to ensure the high quality scanning of the model. This goal is achieved by placing the scanner at strategically selected Next-Best-Views (NBVs) to ensure progressively capturing the geometric details of the object, until both completeness and high fidelity are reached. The technique is based on the analysis of a Poisson field and its geometric relation with an input scan. We generate a confidence map that reflects the quality/fidelity of the estimated Poisson iso-surface. The confidence map guides the generation of a viewing vector field, which is then used for computing a set of NBVs. We applied the algorithm on two different robotic platforms, a PR2 mobile robot and a one-arm industry robot. We demonstrated the advantages of our method through a number of autonomous high quality scannings of complex physical objects, as well as performance comparisons against state-of-the-art methods. Pinxin Long, Hui Huang 0004, Daniel Cohen-Or, Minglun Gong, Oliver Deussen, Baoquan Chen |
ACM Trans. Graph. | 6 |
| 2014 | Morfit: interactive surface reconstruction from incomplete point clouds with curve-driven topology and geometry controlabstractWith significant data missing in a point scan, reconstructing a complete surface with sufficient geometric and topological fidelity is highly challenging. We present an interactive technique for surface reconstruction from incomplete and sparse scans of 3D objects possessing sharp features. A fundamental premise of our interaction paradigm is that directly editing data in 3D is not only counterintuitive but also ineffective, while working with 1D entities (i.e., curves) is a lot more manageable. To this end, we factor 3D editing into two "orthogonal" interactions acting on skeletal and profile curves of the underlying shape, controlling its topology and geometric features, respectively. For surface completion, we introduce a novel skeleton-driven morph-to-fit , or morfit , scheme which reconstructs the shape as an ensemble of generalized cylinders. Morfit is a hybrid operator which optimally interpolates between adjacent curve profiles (the "morph") and snaps the surface to input points (the "fit"). The interactive reconstruction iterates between user edits and morfit to converge to a desired final surface. We demonstrate various interactive reconstructions from point scans with sharp features and significant missing data. Kangxue Yin, Hui Huang 0004, Hao (Richard) Zhang, Minglun Gong, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 4 |
| 2013 | Underwater Camera Calibration Using Wavelength TriangulationabstractIn underwater imagery, the image formation process includes refractions that occur when light passes from water into the camera housing, typically through a flat glass port. We extend the existing work on physical refraction models by considering the dispersion of light, and derive new constraints on the model parameters for use in calibration. This leads to a novel calibration method that achieves improved accuracy compared to existing work. We describe how to construct a novel calibration device for our method and evaluate the accuracy of the method through synthetic and real experiments. Timothy Yau, Minglun Gong, Yee-Hong Yang |
CVPR | 2 |
| 2013 | CIDER: Concept-based image diversification, exploration, and retrieval
Enamul Hoque Prince, Orland Hoeber, Minglun Gong |
Inf. Process. Manag. | 3 |
| 2013 | L1-medial skeleton of point cloudabstractWe introduce L 1 - medial skeleton as a curve skeleton representation for 3D point cloud data. The L 1 -median is well-known as a robust global center of an arbitrary set of points. We make the key observation that adapting L 1 -medians locally to a point set representing a 3D shape gives rise to a one-dimensional structure, which can be seen as a localized center of the shape. The primary advantage of our approach is that it does not place strong requirements on the quality of the input point cloud nor on the geometry or topology of the captured shape. We develop a L 1 -medial skeleton construction algorithm, which can be directly applied to an unoriented raw point scan with significant noise, outliers, and large areas of missing data. We demonstrate L 1 -medial skeletons extracted from raw scans of a variety of shapes, including those modeling high-genus 3D objects, plant-like structures, and curve networks. Hui Huang 0004, Daniel Cohen-Or, Minglun Gong, Hao (Richard) Zhang, Guiqing Li, Baoquan Chen |
ACM Trans. Graph. | 4 |
| 2013 | Edge-aware point set resamplingabstractPoints acquired by laser scanners are not intrinsically equipped with normals, which are essential to surface reconstruction and point set rendering using surfels. Normal estimation is notoriously sensitive to noise. Near sharp features, the computation of noise-free normals becomes even more challenging due to the inherent undersampling problem at edge singularities. As a result, common edge-aware consolidation techniques such as bilateral smoothing may still produce erroneous normals near the edges. We propose a resampling approach to process a noisy and possibly outlier-ridden point set in an edge-aware manner. Our key idea is to first resample away from the edges so that reliable normals can be computed at the samples, and then based on reliable data, we progressively resample the point set while approaching the edge singularities. We demonstrate that our Edge-Aware Resampling (EAR) algorithm is capable of producing consolidated point sets with noise-free normals and clean preservation of sharp features. We also show that EAR leads to improved performance of edge-aware reconstruction methods and point set rendering techniques. Hui Huang 0004, Minglun Gong, Daniel Cohen-Or, Uri M. Ascher, Hao (Richard) Zhang |
ACM Trans. Graph. | 3 |
| 2013 | "Mind the gap": tele-registration for structure-driven image completionabstractConcocting a plausible composition from several non-overlapping image pieces, whose relative positions are not fixed in advance and without having the benefit of priors, can be a daunting task. Here we propose such a method, starting with a set of sloppily pasted image pieces with gaps between them. We first extract salient curves that approach the gaps from non-tangential directions, and use likely correspondences between pairs of such curves to guide a novel tele-registration method that simultaneously aligns all the pieces together. A structure-driven image completion technique is then proposed to fill the gaps, allowing the subsequent employment of standard in-painting tools to finish the job. Hui Huang 0004, Kangxue Yin, Minglun Gong, Dani Lischinski, Daniel Cohen-Or, Uri M. Ascher, Baoquan Chen |
ACM Trans. Graph. | 3 |
| 2013 | Projective analysis for 3D shape segmentationabstractWe introduce projective analysis for semantic segmentation and labeling of 3D shapes. The analysis treats an input 3D shape as a collection of 2D projections, labels each projection by transferring knowledge from existing labeled images, and back-projects and fuses the labelings on the 3D shape. The image-space analysis involves matching projected binary images of 3D objects based on a novel bi-class Hausdorff distance . The distance is topology-aware by accounting for internal holes in the 2D figures and it is applied to piecewise-linearly warped object projections to compensate for part scaling and view discrepancies. Projective analysis simplifies the processing task by working in a lower-dimensional space, circumvents the requirement of having complete and well-modeled 3D shapes, and addresses the data challenge for 3D shape analysis by leveraging the massive available image data. A large and dense labeled set ensures that the labeling of a given projected image can be inferred from closely matched labeled images. We demonstrate semantic labeling of imperfect (e.g., incomplete or self-intersecting) 3D models which would be otherwise difficult to analyze without taking the projective analysis approach. Yunhai Wang, Minglun Gong, Tianhua Wang, Daniel Cohen-Or, Hao (Richard) Zhang, Baoquan Chen |
ACM Trans. Graph. | 2 |
| 2012 | Automatic Real-Time Video Matting Using Time-of-Flight Camera and Multichannel Poisson Equations
Liang Wang 0002, Minglun Gong, Ruigang Yang, Cha Zhang, Yee-Hong Yang |
Int. J. Comput. Vis. | 2 |
| 2012 | Field-guided registration for feature-conforming shape compositionabstractWe present an automatic shape composition method to fuse two shape parts which may not overlap and possibly contain sharp features, a scenario often encountered when modeling man-made objects. At the core of our method is a novel field-guided approach to automatically align two input parts in a feature-conforming manner. The key to our field-guided shape registration is a natural continuation of one part into the ambient field as a means to introduce an overlap with the distant part, which then allows a surface-to-field registration. The ambient vector field we compute is feature-conforming; it characterizes a piecewise smooth field which respects and naturally extrapolates the surface features. Once the two parts are aligned, gap filling is carried out by spline interpolation between matching feature curves followed by piecewise smooth least-squares surface reconstruction. We apply our algorithm to obtain feature-conforming shape composition on a variety of models and demonstrate generality of the method with results on parts with or without overlap and with or without salient features. Hui Huang 0004, Minglun Gong, Daniel Cohen-Or, Yaobin Ouyang, Fuwen Tan, Hao (Richard) Zhang |
ACM Trans. Graph. | 2 |
| 2012 | Video Stereolization: Combining Motion Analysis with User InteractionabstractWe present a semiautomatic system that converts conventional videos into stereoscopic videos by combining motion analysis with user interaction, aiming to transfer as much as possible labeling work from the user to the computer. In addition to the widely used structure from motion (SFM) techniques, we develop two new methods that analyze the optical flow to provide additional qualitative depth constraints. They remove the camera movement restriction imposed by SFM so that general motions can be used in scene depth estimation-the central problem in mono-to-stereo conversion. With these algorithms, the user's labeling task is significantly simplified. We further developed a quadratic programming approach to incorporate both quantitative depth and qualitative depth (such as these from user scribbling) to recover dense depth maps for all frames, from which stereoscopic view can be synthesized. In addition to visual results, we present user study results showing that our approach is more intuitive and less labor intensive, while producing 3D effect comparable to that from current state-of-the-art interactive algorithms. Miao Liao, Jizhou Gao, Ruigang Yang, Minglun Gong |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2011 | Foreground segmentation of live videos using locally competing 1SVMsabstractThe objective of foreground segmentation is to extract the desired foreground object from input videos. Over the years there have been significant amount of efforts on this topic, nevertheless there still lacks a simple yet effective algorithm that can process live videos of objects with fuzzy boundaries captured by freely moving cameras. This paper presents an algorithm toward this goal. The key idea is to train and maintain two competing one-class support vector machines (1SVMs) at each pixel location, which model local color distributions for foreground and background, respectively. We advocate the usage of two competing local classifiers, as it provides higher discriminative power and allows better handling of ambiguities. As a result, our algorithm can deal with a variety of videos with complex backgrounds and freely moving cameras with minimum user interactions. In addition, by introducing novel acceleration techniques and by exploiting the parallel structure of the algorithm, realtime processing speed is achieved for VGA-sized videos. Minglun Gong |
CVPR | 1 |
| 2011 | Data organization and visualization using self-sorting map
Grant Strong, Minglun Gong |
Graphics Interface | 2 |
| 2011 | Incorporating estimated motion in real-time background subtractionabstractMany existing background subtraction approaches model background color only and detect foreground as outliers, and hence may confuse background changes or noises with true foreground. We present a novel algorithm that utilizes motion cues computed from an optical flow algorithm. The additional motion information allows aligning moving foreground objects over time so that models can be built for foreground as well. It also facilities background (and foreground) modeling since both color and motion cues can be utilized. In practice, our GPU implementation is able to process QVGA-sized video sequences at 39.3 FPS on a laptop. Quantitative evaluation on standard testbeds demonstrate the competitive performance of our approach. Minglun Gong, Li Cheng 0001 |
ICIP | 1 |
| 2011 | Similarity-based image organization and browsing using multi-resolution self-organizing map
Grant Strong, Minglun Gong |
Image Vis. Comput. | 2 |
| 2011 | Near-real-time stereo matching with slanted surface modeling and sub-pixel accuracy
Minglun Gong, Yee-Hong Yang |
Pattern Recognit. | 1 |
| 2011 | Real-Time Discriminative Background SubtractionabstractThe authors examine the problem of segmenting foreground objects in live video when background scene textures change over time. In particular, we formulate background subtraction as minimizing a penalized instantaneous risk functional--yielding a local online discriminative algorithm that can quickly adapt to temporal changes. We analyze the algorithm's convergence, discuss its robustness to nonstationarity, and provide an efficient nonlinear extension via sparse kernels. To accommodate interactions among neighboring pixels, a global algorithm is then derived that explicitly distinguishes objects versus background using maximum a posteriori inference in a Markov random field (implemented via graph-cuts). By exploiting the parallel nature of the proposed algorithms, we develop an implementation that can run efficiently on the highly parallel graphics processing unit (GPU). Empirical studies on a wide variety of datasets demonstrate that the proposed approach achieves quality that is comparable to state-of-the-art offline methods, while still being suitable for real-time video analysis ( ≥ 75 fps on a mid-range GPU). Li Cheng 0001, Minglun Gong, Dale Schuurmans, Terry Caelli |
IEEE Trans. Image Process. | 2 |
| 2010 | Real-time video matting using multichannel poisson equations
Minglun Gong, Liang Wang 0002, Ruigang Yang, Yee-Hong Yang |
Graphics Interface | 1 |
| 2009 | Realtime background subtraction from dynamic scenesabstractThis paper examines the problem of moving object detection. More precisely, it addresses the difficult scenarios where background scene textures in the video might change over time. In this paper, we formulate the problem mathematically as minimizing a constrained risk functional motivated from the large margin principle. It is a generalization of the one class support vector machines (1-SVMs) to accommodate spatial interactions, which is further incorporated into an online learning framework to track temporal changes. As a result it yields a closed-form update formula, a central component of the proposed algorithm to enable prompt adaptation to spatio-temporal changes. We also analyze the mistake bound and discuss issues such as dealing with non-stationary distributions, making use of kernels and efficient inference by a variant of dynamic programming. By exploiting the inherently concurrent structure, the proposed approach is designed to work with the highly parallel graphics processors (GPUs) to facilitate realtime analysis. Our empirical study demonstrates that the proposed approach works in realtime (over 80 frames per second) and at the same time performs competitively against state-of-the-art offline and quasi-realtime methods. Li Cheng 0001, Minglun Gong |
ICCV | 2 |
| 2009 | Modeling deformable objects from a single depth cameraabstractWe propose a novel approach to reconstruct complete 3D deformable models over time by a single depth camera, provided that most parts of the models are observed by the camera at least once. The core of this algorithm is based on the assumption that the deformation is continuous and predictable in a short temporal interval. While the camera can only capture part of a whole surface at any time instant, partial surfaces reconstructed from different times are assembled together to form a complete 3D surface for each time instant, even when the shape is under severe deformation. A mesh warping algorithm based on linear mesh deformation is used to align different partial surfaces. A volumetric method is then used to combine partial surfaces, fix missing holes, and smooth alignment errors. Our experiment shows that this approach is able to reconstruct visually plausible 3D surface deformation results with a single camera. Miao Liao, Qing Zhang 0017, Huamin Wang 0001, Ruigang Yang, Minglun Gong |
ICCV | 5 |
| 2009 | Real-time joint disparity and disparity flow estimation on programmable graphics hardware
Minglun Gong |
Comput. Vis. Image Underst. | 1 |
| 2008 | Stereoscopic inpainting: Joint color and depth completion from stereo imagesabstractWe present a novel algorithm for simultaneous color and depth inpainting. The algorithm takes stereo images and estimated disparity maps as input and fills in missing color and depth information introduced by occlusions or object removal. We first complete the disparities for the occlusion regions using a segmentation-based approach. The completed disparities can be used to facilitate the user in labeling objects to be removed. Since part of the removed regions in one image is visible in the other, we mutually complete the two images through 3D warping. Finally, we complete the remaining unknown regions using a depth-assisted texture synthesis technique, which simultaneously fills in both color and depth. We demonstrate the effectiveness of the proposed algorithm on several challenging data sets. Liang Wang 0002, Hailin Jin, Ruigang Yang, Minglun Gong |
CVPR | 4 |
| 2008 | Real-time Light Fall-off StereoabstractWe present a real-time depth recovery system using Light Fall-off Stereo (LFS). Our system contains two co-axial point light sources (LEDs) synchronized with a video camera. The video camera captures the scene under these two LEDs in complementary states(e.g., one on, one off). Based on the inverse square law for light intensity, the depth can be directly solved using the pixel ratio from two consecutive frames. We demonstrate the effectiveness of our approach with a number of real world scenes. Quantitative evaluation shows that our system compares favorably to other commercial real-time 3D range sensors, particularly in textured areas. We believe our system offers a low-cost high-resolution alternative for depth sensing under controlled lighting. Miao Liao, Liang Wang 0002, Ruigang Yang, Minglun Gong |
ICIP | 4 |
| 2008 | Real-time foreground segmentation on GPUs using local online learning and global graph cut optimizationabstractThis paper is to address the problem of foreground separation from the background modeling perspective. In particular, we deal with the difficult scenarios where the background texture might change spatially and temporally. A novel approach is proposed that incorporates a pixel-based online learning method to adapt to temporal background changes promptly, together with a graph cuts method to propagate per-pixel evaluation results over nearby pixels. Empirical experiments on a variety of datasets demonstrate the competitiveness of the proposed approach, which is also able to work in real-time on the Graphics Processing Unit (GPU) of programmable graphics cards. Minglun Gong, Li Cheng 0001 |
ICPR | 1 |
| 2008 | Local stereo matching with 3D adaptive cost aggregation for slanted surface modeling and sub-pixel accuracyabstractThis paper presents a new local binocular stereo algorithm which takes into consideration plane fitting at the per-pixel level. Two disparity calculation passes are used. The first pass assumes that surfaces in the scene are fronto-parallel and generates an initial disparity map, from which the disparity plane orientations of all pixels are extracted and refined. In the second pass, the cost aggregation for each pixel is conducted along the estimated disparity plane orientations, rather than the fronto-parallel ones. Large window size with adaptive support weights is used to ensure the effectiveness of the slanted surface modeling. The disparity search space is also quantized at sub-pixel level to improve the accuracy of the disparity results. The experimental results demonstrate the validity of our presented approach. Minglun Gong, Yee-Hong Yang |
ICPR | 2 |
| 2007 | Light Fall-off StereoabstractWe present light fall-off stereo-LFS-a new method for computing depth from scenes beyond lambertian reflectance and texture. LFS takes a number of images from a stationary camera as the illumination source moves away from the scene. Based on the inverse square law for light intensity, the ratio images are directly related to scene depth from the perspective of the light source. Using this as the invariant, we developed both local and global methods for depth recovery. Compared to previous reconstruction methods for non-lamebrain scenes, LFS needs as few as two images, does not require calibrated camera or light sources, or reference objects in the scene. We demonstrated the effectiveness of LFS with a variety of real-world scenes. Miao Liao, Liang Wang 0002, Ruigang Yang, Minglun Gong |
CVPR | 4 |
| 2007 | Real-time backward disparity-based rendering for dynamic scenes using programmable graphics hardwareabstractThis paper presents a backward disparity-based rendering algorithm, which runs at real-time speed on programmable graphics hardware. The algorithm requires only a handful of image samples of the scene and estimated noisy disparity maps, whereas most existing techniques need either dense samples or accurate depth information. To color a given pixel in the novel view, a backward searching process is conducted to find the corresponding pixels from the closest four reference images. The use of backward searching process makes the algorithm more robust to errors in estimated disparity maps than existing forward warping-based approaches. In addition, since the computations for different pixels are independent, they can be performed in parallel on the Graphics Processing Units of modern graphics hardware. Experiment results demonstrate that our algorithm can synthesize accurate novel views for dynamic real scenes at a high frame rate. Minglun Gong, Jason M. Selzer, Yee-Hong Yang |
Graphics Interface | 1 |
| 2007 | A Performance Study on Different Cost Aggregation Approaches Used in Real-Time Stereo Matching
Minglun Gong, Ruigang Yang, Liang Wang 0002, Mingwei Gong |
Int. J. Comput. Vis. | 1 |
| 2007 | Real-Time Stereo Matching Using Orthogonal Reliability-Based Dynamic ProgrammingabstractA novel algorithm is presented in this paper for estimating reliable stereo matches in real time. Based on the dynamic programming-based technique we previously proposed, the new algorithm can generate semi-dense disparity maps using as few as two dynamic programming passes. The iterative best path tracing process used in traditional dynamic programming is replaced by a local minimum searching process, making the algorithm suitable for parallel execution. Most computations are implemented on programmable graphics hardware, which improves the processing speed and makes real-time estimation possible. The experiments on the four new Middlebury stereo datasets show that, on an ATI Radeon X800 card, the presented algorithm can produce reliable matches for 60% approximately 80% of pixels at the rate of 10 approximately 20 frames per second. If needed, the algorithm can be configured for generating full density disparity maps. Minglun Gong, Yee-Hong Yang |
IEEE Trans. Image Process. | 1 |
| 2006 | Enforcing Temporal Consistency in Real-Time Stereo Estimation
Minglun Gong |
ECCV (3) | 1 |
| 2006 | Estimate Large Motions Using the Reliability-Based Motion Estimation Algorithm
Minglun Gong, Yee-Hong Yang |
Int. J. Comput. Vis. | 1 |
| 2005 | Near Real-Time Reliable Stereo Matching Using Programmable Graphics HardwareabstractA near-real-time stereo matching technique is presented in this paper, which is based on the reliability-based dynamic programming algorithm we proposed earlier. The new algorithm can generate semi-dense disparity maps using only two dynamic programming passes, while our previous approach requires 20-30 passes. We also implement the algorithm on programmable graphics hardware, which further improves the processing speed. The experiments on the four Middlebury stereo datasets show that the new algorithm can produce dense (>85% of the pixels) and reliable (error rate <0.3%) matches in near real-time (0.05-0.1 sec). If needed, it can also be used to generate dense disparity maps. Based on the evaluation conducted by the Middlebury Stereo Vision Research Website, the new algorithm is ranked between the variable window and the graph cuts approaches and currently is the most accurate dynamic programming based technique. When more than one reference images are available, the accuracy can be further improved with little extra computation time. Minglun Gong, Yee-Hong Yang |
CVPR (1) | 1 |
| 2005 | Camera field rendering for static and dynamic scenes
Minglun Gong, Yee-Hong Yang |
Graph. Model. | 1 |
| 2005 | Fast Unambiguous Stereo Matching Using Reliability-Based Dynamic ProgrammingabstractAn efficient unambiguous stereo matching technique is presented in this paper. Our main contribution is to introduce a new reliability measure to dynamic programming approaches in general. For stereo vision application, the reliability of a proposed match on a scanline is defined as the cost difference between the globally best disparity assignment that includes the match and the globally best assignment that does not include the match. A reliability-based dynamic programming algorithm is derived accordingly, which can selectively assign disparities to pixels when the corresponding reliabilities exceed a given threshold. The experimental results show that the new approach can produce dense (> 70 percent of the unoccluded pixels) and reliable (error rate < 0.5 percent) matches efficiently (< 0.2 sec on a 2GHz P4) for the four Middlebury stereo data sets. Minglun Gong, Yee-Hong Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | Estimate large motions using reliability-based dynamic programmingabstractDetecting and estimating motions of fast moving objects has many important applications. However, most existing motion estimation techniques have difficulties in handling large motions in the scene. In this paper, the reliability-based dynamic programming technique proposed by Gong and Yang is extended and applied to large motion estimation problem. Compared with the Gong and Yang approach, the extended algorithm removes the constant penalty assumption and also explicitly enforces the inter-scanline consistency constraint. The experimental results indicate that the new algorithm can effectively estimate velocities for fast moving objects. The algorithm can also be configured to produce sparse but reliable flow fields. Minglun Gong, Yee-Hong Yang |
ICIP | 1 |
| 2004 | Quadtree-based genetic algorithm and its applications to computer vision
Minglun Gong, Yee-Hong Yang |
Pattern Recognit. | 1 |
| 2003 | Fast Stereo Matching Using Reliability-Based Dynamic Programming and Consistency ConstraintsabstractA method for solving binocular and multiview stereo matching problems is presented here. A weak consistency constraint is proposed, which expresses the visibility constraint in the image space. It can be proved that the weak consistency constraint holds for scenes that can be represented by a set of 3D points. As well, also proposed is a new reliability measure for dynamic programming techniques, which evaluates the reliability of a given match. A novel reliability-based dynamic programming algorithm is derived accordingly, which can selectively assign disparity values to pixels when the reliabilities of the corresponding matches exceed a given threshold. Consistency constraints and the new reliability-based dynamic programming algorithm can be combined in an iterative approach. The experimental results show that the iterative approach can produce dense (60-90%) and reliable (total error rate of 0.1-1.1%) matching for binocular stereo datasets. It can also generate promising disparity maps for trinocular and multiview stereo datasets. Minglun Gong, Yee-Hong Yang |
ICCV | 1 |
| 2002 | Genetic-Based Stereo Algorithm and Disparity Map Evaluation
Minglun Gong, Yee-Hong Yang |
Int. J. Comput. Vis. | 1 |
| 2001 | The Rayset and Its Applications
Minglun Gong, Yee-Hong Yang |
Graphics Interface | 1 |
| 2001 | Layer-Based Morphing
Minglun Gong, Yee-Hong Yang |
Graph. Model. | 1 |
| 1997 | A new method for speeding up ray tracing NURBS surfaces
Kaihuai Qin, Minglun Gong, Guan Youjiang |
Comput. Graph. | 2 |
| 1996 | Fast ray tracing NURBS surfaces
Kaihuai Qin, Minglun Gong, Geliang Tong |
J. Comput. Sci. Technol. | 2 |