EDBT 2026 Demo / reviewers in the wild / expert
Xiaopeng Zhang 0001
dblp:45/5193-1
· DBLP profile ↗
170ranked-venue papers
6as first author
72since 2021 · last 2026
0000-0002-0092-6474ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 135 · 5 first-author · 53 since 2021Artificial intelligence and machine learning · 34 · 1 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 9 · 1 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal InteractionabstractDense video captioning jointly localizes and captions salient events in untrimmed videos. Recent methods primarily focus on leveraging additional prior knowledge and advanced multi-task architectures to achieve competitive performance. However, these pipelines rely on implicit modeling that uses frame-level or fragmented video features, failing to capture the temporal coherence across event sequences and comprehensive semantics within visual contexts. To address this, we propose an explicit temporal-semantic modeling framework called Context-Aware Cross-Modal Interaction (CACMI), which leverages both latent temporal characteristics within videos and linguistic semantics from text corpus. Specifically, our model consists of two core components: Cross-modal Frame Aggregation aggregates relevant frames to extract temporally coherent, event-aligned textual features through cross-modal retrieval; and Context-aware Feature Enhancement utilizes query-guided attention to integrate visual dynamics with pseudo-event semantics. Extensive experiments on the ActivityNet Captions and YouCook2 datasets demonstrate that CACMI achieves the state-of-the-art performance on dense video captioning task. Mingda Jia, Weiliang Meng, Zenghuang Fu, Ju Xin, Rongtao Xu, Jiguang Zhang, Xiaopeng Zhang 0001 |
AAAI | 10 |
| 2026 | DehazeGS: Seeing Through Fog with 3D Gaussian SplattingabstractCurrent novel view synthesis methods are typically designed for high-quality and clean input images. However, in foggy scenes, scattering and attenuation can significantly degrade the quality of rendering. Although NeRF-based dehazing approaches have been developed, their reliance on deep fully connected neural networks and per-ray sampling strategies leads to high computational costs. Furthermore, NeRF's implicit representation limits its ability to recover fine-grained details from hazy scenes. To overcome these limitations, we propose DehazeGS, the first physics-driven 3D Gaussian Splatting (3DGS) framework for dehazing. We adopt an explicit Gaussian representation to model fog formation via a physically consistent forward rendering process, enabling reconstruction and rendering of fog-free scenes using only multi-view foggy images as input. Specifically, based on the atmospheric scattering model, we simulate the formation of fog by establishing the transmission function directly on Gaussian primitives via depth-to-transmission mapping. During training, we jointly learn the atmospheric light and scattering coefficients while optimizing the Gaussian representation of foggy scenes. At inference time, we remove the effects of scattering and attenuation in Gaussian distributions and directly render the scene to obtain dehazed views. Experiments on both real-world and synthetic foggy datasets demonstrate that DehazeGS achieves state-of-the-art performance. Yiqun Wang 0001, Aiheng Jiang, Zhengda Lu, Jianwei Guo 0003, Yong Li 0023, Hongxing Qin, Xiaopeng Zhang 0001 |
AAAI | 8 |
| 2026 | Explicit to implicit presentation for 3D unbounded open scenes reconstruction: the survey
Jiguang Zhang, Weiliang Meng, Zhaohui Zhang 0002, Xiaopeng Zhang 0001 |
Expert Syst. Appl. | 5 |
| 2026 | WM-DETR: Dual-branch wavelet-Mamba and sparse attention for robust underwater object detection
Shunpeng Chen, Zenghuang Fu, Longzhao Huang, Shengpeng Xu, Yixian Kong, Changwei Wang 0001, Weiliang Meng, Xiaopeng Zhang 0001 |
Expert Syst. Appl. | 10 |
| 2026 | Adaptive in Adapter: Boosting Open-Vocabulary Semantic Segmentation With Adaptive Dropout AdapterabstractOpen-vocabulary semantic segmentation is a challenging multimedia task that requires segmentation and recognition of unseen word classes during the testing phase. Recent works bridge the gap between closed and open-vocabulary recognition by introducing large-scale visual language models such as CLIP with cross-modal alignment capabilities. To preserve multimodal alignment capabilities, it is common to freeze the parameters of the CLIP and then add additional learnable components such as adapters to expand to downstream tasks. However, for the open-vocabulary semantic segmentation task, the plain adapter suffers from overfitting the closed-vocabulary classes and impairs performance on the open-vocabulary unseen classes. In addition, since CLIP is trained to perform image-level alignment can cause the network to over-focus on partially discriminative regions, resulting in incomplete segmentation masks. To alleviate the above problems, we introduce adaptive dropout adapters to release theAdaptiveInAdapter (i.e.AIA) from the following two aspects:i)A Generalization Feature Selection Adapter (GFSA) is proposed to improve the generalization of network over unseen classes.ii)A Discriminative Region Mask Adapter (DRMA) is proposed for retrofitting CLIP backbone, has provided region free biased features for segmentation mask generation. Meanwhile, our proposed AIA achieves the current state-of-the-art performance on several open-vocabulary semantic segmentation benchmarks. Code is available athttps://github.com/clearxu/AIA. Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Jiguang Zhang, Xiaoqiang Teng, Weiliang Meng, Xiaopeng Zhang 0001 |
IEEE Trans. Multim. | 9 |
| 2026 | Robust detection in complex construction sites: HiPA-DETR with weather-aware and cross-domain generalization
Zenghuang Fu, Muyang Zhang, Changwei Wang 0001, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001 |
Vis. Comput. | 9 |
| 2025 | PanoDiT: Panoramic Videos Generation with Diffusion TransformerabstractAs immersive experiences become increasingly popular, panoramic video has garnered significant attention in both research and applications. The high cost associated with capturing panoramic video underscores the need for efficient prompt-based generation methods. Although recent text-to-video (T2V) diffusion techniques have shown potential in standard video generation, they face challenges when applied to panoramic videos due to substantial differences in content and motion patterns. In this paper, we propose PanoDiT, a framework that utilizes the Diffusion Transformer (DiT) architecture to generate panoramic videos from text descriptions. Unlike traditional methods that rely on UNet-based denoising, our method leverages a transformer architecture for denoising, incorporating both temporal and global attention mechanisms. This ensures coherent frame generation and smooth motion transitions, offering distinct advantages in long-horizon generation tasks. To further enhance motion and consistency in the generated videos, we introduce DTM-LoRA and two panoramic-specific losses. Compared to previous methods, our PanoDiT achieves state-of-the-art performance across various evaluation metrics and user study, with code is available in the supplementary material. Muyang Zhang, Yuzhi Chen, Rongtao Xu, Changwei Wang 0001, Weiliang Meng, Jianwei Guo 0003, Xiaopeng Zhang 0001 |
AAAI | 9 |
| 2025 | Mask-Guided Transformer with Hybrid Supervision for 3D Instance Segmentationabstract3D instance segmentation from point clouds is a classic but strenuous research problem. Recently, Transformer-based methods have dominated 3D instance segmentation, but most previous methods use learnable queries with low instance mask recall. Meanwhile, object queries usually use masked cross-attention involving an iterative optimization process with inaccurate initial instance masks, which may cause the query to fall into sub-optimal situations. In this paper, we propose MHFormer, a novel Mask-guided Transformer with Hybrid supervision for 3D instance segmentation. We initially develop an instance semantic-aware query module to address the issue of low recall. Subsequently, we introduce ground truth masks to proactively refine the prediction masks from the Transformer layer, guiding the query update process. Furthermore, given the scarcity of matched positive samples (e.g., an average of only 13 instances per scene in ScanNet), we introduce a pioneering one-to-many supervision in 3D instance segmentation to enhance matching efficiency. Experiments and ablation studies conducted on ScanNet, ScanNet200, and S3DIS benchmarks validate the efficacy of our approach. Jianwei Guo 0003, Haobo Qin, Yinchang Zhou, Weiliang Meng, Xiaopeng Zhang 0001 |
ICME | 6 |
| 2025 | DiffusionIMU: Diffusion-Based Inertial Navigation with Iterative Motion RefinementabstractInertial navigation enables self-contained localization using only Inertial Measurement Units (IMUs), making it widely applicable in various domains such as navigation, augmented reality, and robotics. However, existing methods suffer from drift accumulation due to the sensor noise and difficulty capturing long-range temporal dependencies, limiting their robustness and accuracy. To address these challenges, we propose DiffusionIMU, a novel diffusion-based framework for inertial navigation. DiffusionIMU enhances direct velocity regression from IMU data through an iterative generative denoising process, progressively refining motion state estimation. It integrates the noise-adaptive feature modulation for sensor variability handling, the feature alignment mechanism for representation consistency, and the diffusion-based temporal modeling to decrease accumulated drift. Experiments show that DiffusionIMU consistently outperforms existing methods, demonstrating superior generalization to unseen users while alleviating the impact of the sensor noise. Xiaoqiang Teng, Shibiao Xu, Zhihao Hao, Deke Guo, Hai-Sheng Li 0002, Weiliang Meng, Xiaopeng Zhang 0001 |
IJCAI | 9 |
| 2025 | AccidentX: A Large-Scale Multimodal BEV Dataset for Traffic Accident Analysis and PreventionabstractWith the rapid development and widespread application of autonomous driving technology, the accurate analysis and prevention of traffic accidents have become critical challenges. However, current traffic accident datasets are often constrained by limited scale and diversity, impeding progress in this field. To address these limitations, we introduce AccidentX, a large-scale multimodal dataset specifically curated for comprehensive traffic accident analysis and prevention. Our AccidentX comprises over 10,000 bird’s-eye view (BEV) videos generated using the CARLA simulator, with detailed annotations covering a wide range of traffic scenarios. In comparison to existing datasets such as nuScenes, our AccidentX offers seven times more video frames and leverages Vision-Language Models (VLMs) and GPT-4o for enhanced scene understanding and decision-making. We also establish a benchmark for state-of-the-art Multimodal Large Language Models (MLLMs) on AccidentX, fostering further research and innovation within the community. AccidentX will be made available as a fully open source resource for the advancement of the autonomous driving safety algorithm community. Muyang Zhang, Mingda Jia, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001 |
IROS | 8 |
| 2025 | Token Masking Transformer for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) is both a promising and challenging task that aims to achieve object localization exclusively through image category labels for supervision. Visual transformers have recently been applied to WSOL, demonstrating significant success through the exploitation of long-range feature dependencies in self-attention mechanisms. However, the transformer-based approach suffers from the same partial activation problem as the CNN-based approach due to the use of the classification task to train self-attention map, i.e., only a few discriminative regions are assigned high attention response and thus the localization map does not cover the whole object. To alleviate this problem, we propose a plug-and-play Token Masking Transformer (TMT) method to help transformer-based WSOL methods to obtain a more complete localization map by dynamic discriminative token masking. Specifically, a batch-wise discriminative token selection strategy is first introduced to flexibly determine the tokens to be masked in each image. Then, we design a token masking transformer block to perform token masking and inspire the network to mine more object-related tokens. Besides, we also design an intermediate token activation loss to further improve the performance of TMT by imposing constraints on intermediate tokens. Extensive experiments demonstrate that our TMT can substantially improve the performance of existing transformer-based methods without increasing the computational cost, and achieves state-of-the-art performance on two mainstream benchmarks. Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Man Zhang 0005, Xiaopeng Zhang 0001 |
IEEE Trans. Multim. | 7 |
| 2025 | ROMOT: Referring-expression-comprehension open-set multi-object tracking
Wei Li 0237, Bowen Li 0014, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001 |
Vis. Comput. | 6 |
| 2025 | Diff-pcg: diffusion point cloud generation conditioned on continuous normalizing flow
Weiliang Meng, Zhongqi Wu, Jianwei Guo 0003, Xiaopeng Zhang 0001 |
Vis. Comput. | 5 |
| 2025 | PDFT: parameter-diminish fine-tuning for transformer-based models
Muyang Zhang, Weiliang Meng, Mingda Jia, Jiaming Gu, Yihua Shao, Changwei Wang 0001, Rongtao Xu, Xiaopeng Zhang 0001 |
Vis. Comput. | 9 |
| 2024 | Spectral Prompt Tuning: Unveiling Unseen Classes for Zero-Shot Semantic SegmentationabstractRecently, CLIP has found practical utility in the domain of pixel-level zero-shot segmentation tasks. The present landscape features two-stage methodologies beset by issues such as intricate pipelines and elevated computational costs. While current one-stage approaches alleviate these concerns and incorporate Visual Prompt Training (VPT) to uphold CLIP's generalization capacity, they still fall short in fully harnessing CLIP's potential for pixel-level unseen class demarcation and precise pixel predictions. To further stimulate CLIP's zero-shot dense prediction capability, we propose SPT-SEG, a one-stage approach that improves CLIP's adaptability from image to pixel. Specifically, we initially introduce Spectral Prompt Tuning (SPT), incorporating spectral prompts into the CLIP visual encoder's shallow layers to capture structural intricacies of images, thereby enhancing comprehension of unseen classes. Subsequently, we introduce the Spectral Guided Decoder (SGD), utilizing both high and low-frequency information to steer the network's spatial focus towards more prominent classification features, enabling precise pixel-level prediction outcomes. Through extensive experiments on two public datasets, we demonstrate the superiority of our method over state-of-the-art approaches, performing well across all classes and particularly excelling in handling unseen classes. Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Li Guo 0004, Man Zhang 0005, Xiaopeng Zhang 0001 |
AAAI | 7 |
| 2024 | SARNet: Semantic Augmented Registration of Large-Scale Urban Point Clouds
Haobo Qin, Yinchang Zhou, Xiaopeng Zhang 0001, Zhanglin Cheng, Jianwei Guo 0003 |
CVM (1) | 4 |
| 2024 | SVDTree: Semantic Voxel Diffusion for Single Image Tree ReconstructionabstractEfficiently representing and reconstructing the 3D geometry of biological trees remains a challenging problem in computer vision and graphics. We propose a novel approach for generating realistic tree models from single-view photographs. We cast the 3D information inference problem to a semantic voxel diffusion process, which converts an input image of a tree to a novel Semantic Voxel Structure (SVS) in 3D space. The SVS encodes the geometric appearance and semantic structural information (e.g., classifying trunks, branches, and leaves), which retains the intricate internal tree features. Tailored to the SVS, we present SVDTree a new hybrid tree modeling approach by combining structure-oriented branch reconstruction and self-organization-based foliage reconstruction. We validate SVDTree by using images from both synthetic and real trees. The comparison results show that our approach can better preserve tree details and achieve more realistic and accurate reconstruction results than previous methods. Bedrich Benes, Xiaopeng Zhang 0001, Jianwei Guo 0003 |
CVPR | 4 |
| 2024 | UnionFormer: Unified-Learning Transformer with Multi-View Representation for Image Manipulation Detection and LocalizationabstractWe present UnionFormer, a novel framework that inte-grates tampering clues across three views by unified learning for image manipulation detection and localization. Specifically, we construct a BSFI-Net to extract tampering features from RGB and noise views, achieving enhanced responsive-ness to boundary artifacts while modulating spatial consis-tency at different scales. Additionally, to explore the incon-sistency between objects as a new view of clues, we combine object consistency modeling with tampering detection and localization into a three-task unified learning process, allowing them to promote and improve mutually. Therefore, we acquire a unified manipulation discriminative representation under multi-scale supervision that consolidates information from three views. This integration facilitates highly effective concurrent detection and localization of tampering. We perform extensive experiments on diverse datasets, and the results show that the proposed approach outperforms state-of-the-art methods in tampering detection and localization. Shuaibo Li, Wei Ma 0008, Jianwei Guo 0003, Shibiao Xu, Benchong Li, Xiaopeng Zhang 0001 |
CVPR | 6 |
| 2024 | DefFusion: Deformable Multimodal Representation Fusion for 3D Semantic SegmentationabstractThe complementarity between camera and LiDAR data makes fusion methods a promising approach to improve 3D semantic segmentation performance. Recent transformer-based methods have also demonstrated superiority in segmentation. However, multimodal solutions incorporating transformers are underexplored and face two key inherent difficulties: over-attention and noise from different modal data. To overcome these challenges, we propose a Deformable Multimodal Representation Fusion (DefFusion) framework consisting mainly of a Deformable Representation Fusion Transformer and Dynamic Representation Augmentation Modules. The Deformable Representation Fusion Transformer introduces the deformable mechanism in multimodal fusion, avoiding over-attention and improving efficiency by adaptively modeling a 2D key/value set for a given 3D query, thus enabling multimodal fusion with higher flexibility. To enhance the 2D representation and 3D representation, the Dynamic Representation Enhancement Module is proposed to dynamically remove noise in the input representation via Dynamic Grouped Representation Generation and Dynamic Mask Generation. Extensive experiments validate that our model achieves the best 3D semantic segmentation performance on SemanticKITTI and NuScenes benchmarks. Rongtao Xu, Changwei Wang 0001, Duzhen Zhang, Man Zhang 0005, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001 |
ICRA | 7 |
| 2024 | DRC-NET: Density Reweighted Convolution Network for Edge Curve Extraction
Xiaojuan Ning, Qishuai Shi, Yuexuan Liu, Haiyan Jin, Yinghui Wang 0001, Xiaopeng Zhang 0001, Jianwei Guo 0003 |
PRCV (2) | 6 |
| 2024 | InstanceTex: Instance-level Controllable Texture Synthesis for 3D Scenes via Diffusion Priors
Mingxin Yang, Jianwei Guo 0003, Yuzhi Chen, Zhanglin Cheng, Xiaopeng Zhang 0001, Hui Huang 0004 |
SIGGRAPH Asia | 7 |
| 2024 | FEKNN: A Wi-Fi Indoor Localization Method Based on Feature Enhancement and KNN
Bowen Li 0014, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001 |
WASA (1) | 6 |
| 2024 | Self-supervised reconstruction of re-renderable facial textures from single image
Mingxin Yang, Jianwei Guo 0003, Xiaopeng Zhang 0001, Zhanglin Cheng |
Comput. Graph. | 3 |
| 2024 | De-NeRF: Ultra-high-definition NeRF with deformable net alignmentabstractAbstract Neural Radiance Field (NeRF) can render complex 3D scenes with viewpoint‐dependent effects. However, less work has been devoted to exploring its limitations in high‐resolution environments, especially when upscaled to ultra‐high resolution (e.g., 4k). Specifically, existing NeRF‐based methods face severe limitations in reconstructing high‐resolution real scenes, for example, a large number of parameters, misalignment of the input data, and over‐smoothing of details. In this paper, we present a novel and effective framework, called De‐NeRF, based on NeRF and deformable convolutional network, to achieve high‐fidelity view synthesis in ultra‐high resolution scenes: (1) marrying the deformable convolution unit which can solve the problem of misaligned input of the high‐resolution data. (2) Presenting a density sparse voxel‐based approach which can greatly reduce the training time while rendering results with higher accuracy. Compared to existing high‐resolution NeRF methods, our approach improves the rendering quality of high‐frequency details and achieves better visual effects in 4K high‐resolution scenes. Jianing Hou, Runjie Zhang, Zhongqi Wu, Weiliang Meng, Xiaopeng Zhang 0001, Jianwei Guo 0003 |
Comput. Animat. Virtual Worlds | 5 |
| 2024 | SocialVis: Dynamic social visualization in dense scenes via real-time multi-object tracking and proximity graph constructionabstractAbstract To monitor and assess social dynamics and risks at large gatherings, we propose “SocialVis,” a comprehensive monitoring system based on multi‐object tracking and graph analysis techniques. Our SocialVis includes a camera detection system that operates in two modes: a real‐time mode, which enables participants to track and identify close contacts instantly, and an offline mode that allows for more comprehensive post‐event analysis. The dual functionality not only aids in preventing mass gatherings or overcrowding by enabling the issuance of alerts and recommendations to organizers, but also allows for the generation of proximity‐based graphs that map participant interactions, thereby enhancing the understanding of social dynamics and identifying potential high‐risk areas. It also provides tools for analyzing pedestrian flow statistics and visualizing paths, offering valuable insights into crowd density and interaction patterns. To enhance system performance, we designed the SocialDetect algorithm in conjunction with the BYTE tracking algorithm. This combination is specifically engineered to improve detection accuracy and minimize ID switches among tracked objects, leveraging the strengths of both algorithms. Experiments on both public and real‐world datasets validate that our SocialVis outperforms existing methods, showing improvement in detection accuracy and reduction in ID switches in dense pedestrian scenarios. Bowen Li 0014, Wei Li 0237, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001 |
Comput. Animat. Virtual Worlds | 6 |
| 2024 | HIDE: Hierarchical iterative decoding enhancement for multi-view 3D human parameter regressionabstractAbstract Parametric human modeling are limited to either single‐view frameworks or simple multi‐view frameworks, failing to fully leverage the advantages of easily trainable single‐view networks and the occlusion‐resistant capabilities of multi‐view images. The prevalent presence of object occlusion and self‐occlusion in real‐world scenarios leads to issues of robustness and accuracy in predicting human body parameters. Additionally, many methods overlook the spatial connectivity of human joints in the global estimation of model pose parameters, resulting in cumulative errors in continuous joint parameters.To address these challenges, we propose a flexible and efficient iterative decoding strategy. By extending from single‐view images to multi‐view video inputs, we achieve local‐to‐global optimization. We utilize attention mechanisms to capture the rotational dependencies between any node in the human body and all its ancestor nodes, thereby enhancing pose decoding capability. We employ a parameter‐level iterative fusion of multi‐view image data to achieve flexible integration of global pose information, rapidly obtaining appropriate projection features from different viewpoints, ultimately resulting in precise parameter estimation. Through experiments, we validate the effectiveness of the HIDE method on the Human3.6M and 3DPW datasets, demonstrating significantly improved visualization results compared to previous methods. Weitao Lin, Jiguang Zhang, Weiliang Meng, Xianglong Liu 0007, Xiaopeng Zhang 0001 |
Comput. Animat. Virtual Worlds | 5 |
| 2024 | Key-point-guided adaptive convolution and instance normalization for continuous transitive face reenactment of any personabstractAbstract Face reenactment technology is widely applied in various applications. However, the reconstruction effects of existing methods are often not quite realistic enough. Thus, this paper proposes a progressive face reenactment method. First, to make full use of the key information, we propose adaptive convolution and instance normalization to encode the key information into all learnable parameters in the network, including the weights of the convolution kernels and the means and variances in the normalization layer. Second, we present continuous transitive facial expression generation according to all the weights of the network generated by the key points, resulting in the continuous change of the image generated by the network. Third, in contrast to classical convolution, we apply the combination of depth‐ and point‐wise convolutions, which can greatly reduce the number of weights and improve the efficiency of training. Finally, we extend the proposed face reenactment method to the face editing application. Comprehensive experiments demonstrate the effectiveness of the proposed method, which can generate a clearer and more realistic face from any person and is more generic and applicable than other methods. Shibiao Xu, Miao Hua, Jiguang Zhang, Zhaohui Zhang 0002, Xiaopeng Zhang 0001 |
Comput. Animat. Virtual Worlds | 5 |
| 2024 | AG-SDM: Aquascape generation based on stable diffusion model with low-rank adaptationabstractAbstract As an amalgamation of landscape design and ichthyology, aquascape endeavors to create visually captivating aquatic environments imbued with artistic allure. Traditional methodologies in aquascape, governed by rigid principles such as composition and color coordination, may inadvertently curtail the aesthetic potential of the landscapes. In this paper, we propose Aquascape Generation based on Stable Diffusion Models (AG‐SDM), prioritizing aesthetic principles and color coordination to offer guiding principles for real artists in Aquascape creation. We meticulously curated and annotated three aquascape datasets with varying aspect ratios to accommodate diverse landscape design requirements regarding dimensions and proportions. Leveraging the Fréchet Inception Distance (FID) metric, we trained AGFID for quality assessment. Extensive experiments validate that our AG‐SDM excels in generating hyper‐realistic underwater landscape images, closely resembling real flora, and achieves state‐of‐the‐art performance in aquascape image generation. Muyang Zhang, Yuewei Xian, Wei Li 0237, Jiaming Gu, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001 |
Comput. Animat. Virtual Worlds | 8 |
| 2024 | DomainFeat: Learning Local Features With Domain AdaptationabstractAccurate and efficient keypoint detection and description is a fundamental step in various computer vision tasks. In this paper, we extract robust descriptors and detect accurate keypoints by learning local Features with Domain adaptation (DomainFeat). Specifically, our Domainfeat includes image-level domain invariance supervision, pixel-level domain consistency supervision, Pixel-Adaptive keypoint Detection(PA-Det), and cross-domain dataset with domain stable point supervision. First, we introduce the image-level domain invariance supervision to make the high-level feature distributions from different domains close by fusing domain-invariant representations in the decoder. Furthermore, to compensate for the inconsistency between descriptors corresponding to the keypoints at the pixel level, we propose the pixel-level domain consistency supervision. Then we present the Pixel-Adaptive keypoint Detection to efficiently detect accurate keypoints, which can improve accuracy by enhancing the local consistency of heatmaps. Finally, we propose an efficient approach to construct data and supervision labels in diverse domains, which can tackle complex application scenarios. With these novel modules and supervision methods, our DomainFeat can make feature detectors more accurate and descriptors more robust. Extensive experiments confirm that Domainfeat achieves state-of-the-art performance on benchmarks such as Aachen-Day-Night localization, HPatches image matching, and the challenging DNIM dataset. Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Bin Fan 0001, Xiaopeng Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Exploring Intrinsic Discrimination and Consistency for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) is a challenging and promising task that aims to localize objects solely based on the supervision of image category labels. In the absence of annotated bounding boxes, WSOL methods must employ the intrinsic properties of the image classification task pipeline to generate object localizations. In this work, we propose a WSOL method for exploring the Intrinsic Discrimination and Consistency in the image classification task pipeline, and call it as IDC. First, we develop a Triplet Metrics Based Foreground Modeling (TMFM) framework to directly predict object foreground regions using intrinsic discrimination. Unlike Class Activation Map (CAM) based methods that also rely on intrinsic discrimination, our TMFM framework alleviates the problem of only focusing on the most discriminative parts by optimizing foreground and background regions synergistically. Second, we design a Dual Geometric Transformation Consistency Constraints (DGTC2) training strategy to introduce additional supervision and regularization constraints for WSOL by leveraging intrinsic geometric transformation consistency. The proposed pixel-wise and object-wise consistency constraint losses cost-effectively provide spontaneous supervision for WSOL. Extensive experiments show that our IDC method achieves significant and consistent performance gains compared to existing state-of-the-art WSOL approaches. Code is available at: https://github.com/vignywang/IDC. Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Ruisheng Wang 0001, Xiaopeng Zhang 0001 |
IEEE Trans. Image Process. | 6 |
| 2024 | SkinFormer: Learning Statistical Texture Representation With Transformer for Skin Lesion SegmentationabstractAccurate skin lesion segmentation from dermoscopic images is of great importance for skin cancer diagnosis. However, automatic segmentation of melanoma remains a challenging task because it is difficult to incorporate useful texture representations into the learning process. Texture representations are not only related to the local structural information learned by CNN, but also include the global statistical texture information of the input image. In this paper, we propose a transFormer network (SkinFormer) that efficiently extracts and fuses statistical texture representation for Skin lesion segmentation. Specifically, to quantify the statistical texture of input features, a Kurtosis-guided Statistical Counting Operator is designed. We propose Statistical Texture Fusion Transformer and Statistical Texture Enhance Transformer with the help of Kurtosis-guided Statistical Counting Operator by utilizing the transformer's global attention mechanism. The former fuses structural texture information and statistical texture information, and the latter enhances the statistical texture of multi-scale features. Extensive experiments on three publicly available skin lesion datasets validate that our SkinFormer outperforms other SOAT methods, and our method achieves 93.2% Dice score on ISIC 2018. It can be easy to extend SkinFormer to segment 3D images in the future. Rongtao Xu, Changwei Wang 0001, Jiguang Zhang, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | DTTCNet: Time-to-Collision Estimation With Autonomous Emergency Braking Using Multi-Scale Transformer NetworkabstractThe rapid advancement of autonomous driving technologies has brought the significance of Autonomous Emergency Braking (AEB) systems, which are paramount in mitigating collision risk and elevating road safety by preemptively applying brakes when a potential collision is detected. Within the core mechanisms of AEB systems, the Time-to-Collision (TTC) estimation plays a pivotal role, in quantitatively determining the criticality and timing for initiating braking interventions. However, existing TTC estimation approaches exhibit sensitivity to diverse driving scenarios, compromising the performance of AEB systems, especially in instantaneous situations. To address these issues, this paper presents DTTCNet, a novel supervised deep learning model for TTC estimation that leverages multi-scale transformer architectures and multi-task losses, thereby enhancing precision and boosting system performance. The DTTCNet first extracts spatiotemporal features from raw sensor data and utilizes a supervised training strategy. The multi-scale transformer architecture effectively captures variations across different scales, while the multi-task loss function optimizes the network training performance. Our experimental results on a challenging dataset demonstrate that DTTCNet achieves approximately 20% performance improvements over existing methods in terms of accuracy. This signifies a promising approach to augmenting the safety of autonomous driving systems with the integration of aftermarket mobile devices (e.g., Mobileye and Bosch products). Xiaoqiang Teng, Shibiao Xu, Deke Guo, Yulan Guo, Weiliang Meng, Xiaopeng Zhang 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | Wave-Like Class Activation Map With Representation Fusion for Weakly-Supervised Semantic SegmentationabstractThe Class Activation Map (CAM) is widely used to generate pseudo-labels for Weakly Supervised Semantic Segmentation (WSSS), while it does not adequately consider the modeling of foreground-independent information, resulting in prone to false positive pixels. In this paper, we propose a Wave-like Class Activation Map (WaveCAM) from the perspective of representation fusion and dynamic aggregation representation to alleviate the above problem. Specifically, our WaveCAM includes the foreground-aware representation modeling that enhances perception of foreground information, and the foreground-independent representation modeling that enhances perception of foreground-independent information, and a representation-adaptive fusion module that fuses the two representations. Both representations are expressed as wave functions with amplitude and phase to dynamically aggregate representations and extract semantic information after initialization, and they are fused through the adaptive fusion module to obtain an output containing rich semantic information. Extensive experiments on PASCAL VOC 2012 dataset and MS COCO 2014 dataset validate that our WaveCAM can easily embed multi-stage WSSS and end-to-end WSSS, achieving the state-of-the-art performance. Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Accurate Lung Nodule Segmentation With Detailed Representation Transfer and Soft Mask SupervisionabstractAccurate lung lesion segmentation from computed tomography (CT) images is crucial to the analysis and diagnosis of lung diseases, such as COVID-19 and lung cancer. However, the smallness and variety of lung nodules and the lack of high-quality labeling make the accurate lung nodule segmentation difficult. To address these issues, we first introduce a novel segmentation mask named " soft mask," which has richer and more accurate edge details description and better visualization, and develop a universal automatic soft mask annotation pipeline to deal with different datasets correspondingly. Then, a novel network with detailed representation transfer and soft mask supervision (DSNet) is proposed to process the input low-resolution images of lung nodules into high-quality segmentation results. Our DSNet contains a special detailed representation transfer module (DRTM) for reconstructing the detailed representation to alleviate the small size of lung nodules images and an adversarial training framework with soft mask for further improving the accuracy of segmentation. Extensive experiments validate that our DSNet outperforms other state-of-the-art methods for accurate lung nodule segmentation, and has strong generalization ability in other accurate medical segmentation tasks with competitive results. Besides, we provide a new challenging lung nodules segmentation dataset for further studies (https://drive.google.com/file/d/15NNkvDTb_0Ku0IoPsNMHezJRTH1Oi1wm/view?usp=sharing). Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Jun Xiao 0005, Xiaopeng Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Line-Based 3D Building Abstraction and Polygonal Surface Reconstruction From ImagesabstractTextureless objects, repetitive patterns and limited computational resources pose significant challenges to man-made structure reconstruction from images, because feature-points-based reconstruction methods usually fail due to the lack of distinct texture or ambiguous point matches. Meanwhile multi-view stereo approaches also suffer from high computational complexity. In this article, we present a new framework to reconstruct 3D surfaces for buildings from multi-view images by leveraging another fundamental geometric primitive: line segments. To this end, we first propose a new multi-resolution line segment detector to extract 2D line segments from each image. Then, we construct a 3D line cloud by introducing an improved Line3D++ algorithm to match 2D line segments from different images. Finally, we reconstruct a complete and manifold surface mesh from 3D line segments by formulating a Bayesian probabilistic modeling problem, which accurately generates a set of underlying planes. This output model is simple and has low performance requirements for hardware devices. Experimental results demonstrate the validity of the proposed approach and its ability to generate abstract and compact surface meshes from the 3D line cloud with low computational costs. Jianwei Guo 0003, Xiaopeng Zhang 0001, Zhanglin Cheng |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | Self Correspondence Distillation for End-to-End Weakly-Supervised Semantic SegmentationabstractEfficiently training accurate deep models for weakly supervised semantic segmentation (WSSS) with image-level labels is challenging and important. Recently, end-to-end WSSS methods have become the focus of research due to their high training efficiency. However, current methods suffer from insufficient extraction of comprehensive semantic information, resulting in low-quality pseudo-labels and sub-optimal solutions for end-to-end WSSS. To this end, we propose a simple and novel Self Correspondence Distillation (SCD) method to refine pseudo-labels without introducing external supervision. Our SCD enables the network to utilize feature correspondence derived from itself as a distillation target, which can enhance the network's feature learning process by complementing semantic information. In addition, to further improve the segmentation accuracy, we design a Variation-aware Refine Module to enhance the local consistency of pseudo-labels by computing pixel-level variation. Finally, we present an efficient end-to-end Transformer-based framework (TSCD) via SCD and Variation-aware Refine Module for the accurate WSSS task. Extensive experiments on the PASCAL VOC 2012 and MS COCO 2014 datasets demonstrate that our method significantly outperforms other state-of-the-art methods. Our code is available at https://github.com/Rongtao-Xu/RepresentationLearning/tree/main/SCD-AAAI2023. Rongtao Xu, Changwei Wang 0001, Jiaxi Sun, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001 |
AAAI | 6 |
| 2023 | Audio-Driven Lips and Expression on 3D Human Face
Weiliang Meng, Shibiao Xu, Xiaopeng Zhang 0001 |
CGI | 5 |
| 2023 | SECAD-Net: Self-Supervised CAD Reconstruction by Learning Sketch-Extrude OperationsabstractReverse engineering CAD models from raw geometry is a classic but strenuous research problem. Previous learning-based methods rely heavily on labels due to the supervised design patterns or reconstruct CAD shapes that are not easily editable. In this work, we introduce SECADNet, an end-to-end neural network aimed at reconstructing compact and easy-to-edit CAD models in a self-supervised manner. Drawing inspiration from the modeling language that is most commonly used in modern CAD software, we propose to learn 2D sketches and 3D extrusion parameters from raw shapes, from which a set of extrusion cylinders can be generated by extruding each sketch from a 2D plane into a 3D body. By incorporating the Boolean operation (i.e., union), these cylinders can be combined to closely approximate the target geometry. We advocate the use of implicit fields for sketch representation, which allows for creating CAD variations by interpolating latent codes in the sketch latent space. Extensive experiments on both ABC and Fusion 360 datasets demonstrate the effectiveness of our method, and show superiority over state-of-the-art alternatives including the closely related method for supervised CAD reconstruction. We further apply our approach to CAD editing and single-view CAD reconstruction. Code will be released at https://github.com/BunnySoCrazy/SECAD-Net. Jianwei Guo 0003, Xiaopeng Zhang 0001, Dong-Ming Yan 0001 |
CVPR | 3 |
| 2023 | Treating Pseudo-labels Generation as Image Matting for Weakly Supervised Semantic SegmentationabstractGenerating accurate pseudo-labels under the supervision of image categories is a crucial step in Weakly Supervised Semantic Segmentation (WSSS). In this work, we propose a Mat-Label pipeline that provides a fresh way to treat WSSS pseudo-labels generation as an image matting task. By taking a trimap as input which specifies the foreground, background and unknown regions, the image matting task outputs an object mask with fine edges. The intuition behind our Mat-Label is that generating trimap is much easier than generating pseudo-labels directly under weakly supervised setting. Although current CAM-based methods are off-the-shelf solutions for generating a trimap, they suffer from cross-category and foreground-background pixel prediction confusion. To solve this problem, we develop a Double Decoupled Class Activation Map (D2CAM) for Mat-Label to generate a high-quality trimap. By drawing on the idea of metric learning, we explicitly model class activation map with category decoupling and foreground-background decoupling. We also design two simple yet effective refinement constraints for D2CAM to stabilize optimization and eliminate non-exclusive activation. Extensive experiments validate that our Mat-Label achieves substantial and consistent performance gains compared to current state-of-the-art WSSS approaches. Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001 |
ICCV | 5 |
| 2023 | FeaCo: Reaching Robust Feature-Level Consensus in Noisy Pose ConditionsabstractCollaborative perception offers a promising solution to overcome challenges such as occlusion and long-range data processing. However, limited sensor accuracy leads to noisy poses that misalign observations among vehicles. To address this problem, we propose the FeaCo, which achieves robust Feature-level Consensus among collaborating agents in noisy pose conditions without additional training. We design an efficient Pose-error Rectification Module (PRM) to align derived feature maps from different vehicles, reducing the adverse effect of noisy pose and bandwidth requirements. We also provide an effective multi-scale Cross-level Attention Module (CAM) to enhance information aggregation and interaction between various scales. Our FeaCo outperforms all other localization rectification methods, as validated on both the collaborative perception simulation dataset OPV2V and real-world dataset V2V4Real, reducing heading error and enhancing localization accuracy across various error levels. Our code is available at: https://github.com/jmgu0212/FeaCo.git. Jiaming Gu, Muyang Zhang, Weiliang Meng, Shibiao Xu, Jiguang Zhang, Xiaopeng Zhang 0001 |
ACM Multimedia | 7 |
| 2023 | Deep Deformation Detail Synthesis for Thin Shell ModelsabstractAbstract In physics‐based cloth animation, rich folds and detailed wrinkles are achieved at the cost of expensive computational resources and huge labor tuning. Data‐driven techniques make efforts to reduce the computation significantly by utilizing a preprocessed database. One type of methods relies on human poses to synthesize fitted garments, but these methods cannot be applied to general cloth animations. Another type of methods adds details to the coarse meshes obtained through simulation, which does not have such restrictions. However, existing works usually utilize coordinate‐based representations which cannot cope with large‐scale deformation, and requires dense vertex correspondences between coarse and fine meshes. Moreover, as such methods only add details, they require coarse meshes to be sufficiently close to fine meshes, which can be either impossible, or require unrealistic constraints to be applied when generating fine meshes. To address these challenges, we develop a temporally and spatially as‐consistent‐as‐possible deformation representation (named TS‐ACAP) and design a DeformTransformer network to learn the mapping from low‐resolution meshes to ones with fine details. This TS‐ACAP representation is designed to ensure both spatial and temporal consistency for sequential large‐scale deformations from cloth animations. With this TS‐ACAP representation, our DeformTransformer network first utilizes two mesh‐based encoders to extract the coarse and fine features using shared convolutional kernels, respectively. To transduct the coarse features to the fine ones, we leverage the spatial and temporal Transformer network that consists of vertex‐level and frame‐level attention mechanisms to ensure detail enhancement and temporal coherence of the prediction. Experimental results show that our method is able to produce reliable and realistic animations in various datasets at high frame rates with superior detail synthesis abilities compared to existing methods. Lin Gao 0004, Jie Yang 0038, Shibiao Xu, Juntao Ye, Xiaopeng Zhang 0001, Yukun Lai |
Comput. Graph. Forum | 6 |
| 2023 | Joint specular highlight detection and removal in single images via Unet-TransformerabstractSpecular highlight detection and removal is a fundamental problem in computer vision and image processing. In this paper, we present an efficient end-to-end deep learning model for automatically detecting and removing specular highlights in a single image. In particular, an encoder—decoder network is utilized to detect specular highlights, and then a novel Unet-Transformer network performs highlight removal; we append transformer modules instead of feature maps in the Unet architecture. We also introduce a highlight detection module as a mask to guide the removal task. Thus, these two networks can be jointly trained in an effective manner. Thanks to the hierarchical and global properties of the transformer mechanism, our framework is able to establish relationships between continuous self-attention layers, making it possible to directly model the mapping between the diffuse area and the specular highlight area, and reduce indeterminacy within areas containing strong specular highlight reflection. Experiments on public benchmark and real-world images demonstrate that our approach outperforms state-of-the-art methods for both highlight detection and removal tasks. Zhongqi Wu, Jianwei Guo 0003, Chuanqing Zhuang, Jun Xiao 0005, Dong-Ming Yan 0001, Xiaopeng Zhang 0001 |
Comput. Vis. Media | 6 |
| 2023 | Automatic polyp segmentation via image-level and surrounding-level context fusion deep neural network
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | Dual-stream Representation Fusion Learning for accurate medical image segmentation
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | PuzzleNet: Boundary-Aware Feature Matching for Non-Overlapping 3D Point Clouds Assembly
Jianwei Guo 0003, Haiyong Jiang, Yan-Chao Liu, Xiaopeng Zhang 0001, Dong-Ming Yan 0001 |
J. Comput. Sci. Technol. | 5 |
| 2023 | RC-Net: Row and Column Network with Text Feature for Parsing Floor Plan Images
Weiliang Meng, Zhengda Lu, Jianwei Guo 0003, Jun Xiao 0005, Xiaopeng Zhang 0001 |
J. Comput. Sci. Technol. | 6 |
| 2023 | Attention Weighted Local DescriptorsabstractLocal features detection and description are widely used in many vision applications with high industrial and commercial demands. With large-scale applications, these tasks raise high expectations for both the accuracy and speed of local features. Most existing studies on local features learning focus on the local descriptions of individual keypoints, which neglect their relationships established from global spatial awareness. In this paper, we present AWDesc with a consistent attention mechanism (CoAM) that opens up the possibility for local descriptors to embrace image-level spatial awareness in both the training and matching stages. For local features detection, we adopt local features detection with feature pyramid to obtain more stable and accurate keypoints localization. For local features description, we provide two versions of AWDesc to cope with different accuracy and speed requirements. On the one hand, we introduce Context Augmentation to address the inherent locality of convolutional neural networks by injecting non-local context information, so that local descriptors can "look wider to describe better". Specifically, well-designed Adaptive Global Context Augmented Module (AGCA) and Diverse Surrounding Context Augmented Module (DSCA) are proposed to construct robust local descriptors with context information from global to surrounding. On the other hand, we design an extremely lightweight backbone network coupled with the proposed special knowledge distillation strategy to achieve the best trade-off in accuracy and speed. What is more, we perform thorough experiments on image matching, homography estimation, visual localization, and 3D reconstruction tasks, and the results demonstrate that our method surpasses the current state-of-the-art local descriptors. Code is available at: https://github.com/vignywang/AWDesc. Changwei Wang 0001, Rongtao Xu, Ke Lu 0002, Shibiao Xu, Weiliang Meng, Bin Fan 0001, Xiaopeng Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | Toward Accurate and Efficient Road Extraction by Leveraging the Characteristics of Road ShapesabstractAutomatically extracting roads from very high resolution (VHR) remote sensing images is of great importance in a wide range of remote sensing applications. However, complex shapes of roads (i.e., long, geometrically deformed, and thin) always affected the extraction accuracy, which is one of the challenges of road extraction. Based on the insight into road shape characteristics, we propose a novel road shape aware network (RSANet) to achieve efficient and accurate road extraction. First, we introduce the Efficient Strip Transformer Module (ESTM) to efficiently capture the global context to model the long-distance dependence required by the long roads. Second, we design a Geometric Deformation Estimation Module (GDEM) to adaptively extract the context from the shape deformation caused by shooting roads from different perspectives. Third, we provide a simple but effective Road Edge Focal Loss (REF loss) to make the network focus on optimizing the pixels around the road to alleviate the unbalanced distribution of foreground and background pixels caused by the roads being too thin. Finally, we conduct extensive evaluations on public datasets to verify the effectiveness of RSANet and each of the proposed components. Experiments validate that our RSANet outperforms state-of-the-art methods for road extraction in remote sensing images. Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Ruisheng Wang 0001, Jiguang Zhang, Xiaopeng Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | RSSFormer: Foreground Saliency Enhancement for Remote Sensing Land-Cover SegmentationabstractHigh spatial resolution (HSR) remote sensing images contain complex foreground-background relationships, which makes the remote sensing land cover segmentation a special semantic segmentation task. The main challenges come from the large-scale variation, complex background samples and imbalanced foreground-background distribution. These issues make recent context modeling methods sub-optimal due to the lack of foreground saliency modeling. To handle these problems, we propose a Remote Sensing Segmentation framework (RSSFormer), including Adaptive TransFormer Fusion Module, Detail-aware Attention Layer and Foreground Saliency Guided Loss. Specifically, from the perspective of relation-based foreground saliency modeling, our Adaptive Transformer Fusion Module can adaptively suppress background noise and enhance object saliency when fusing multi-scale features. Then our Detail-aware Attention Layer extracts the detail and foreground-related information via the interplay of spatial attention and channel attention, which further enhances the foreground saliency. From the perspective of optimization-based foreground saliency modeling, our Foreground Saliency Guided Loss can guide the network to focus on hard samples with low foreground saliency responses to achieve balanced optimization. Experimental results on LoveDA datasets, Vaihingen datasets, Potsdam datasets and iSAID datasets validate that our method outperforms existing general semantic segmentation methods and remote sensing segmentation methods, and achieves a good compromise between computational overhead and accuracy. Our code is available at https://github.com/Rongtao-Xu/RepresentationLearning/tree/main/RSSFormer-TIP2023. Rongtao Xu, Changwei Wang 0001, Jiguang Zhang, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001 |
IEEE Trans. Image Process. | 6 |
| 2023 | CNDesc: Cross Normalization for Local Descriptors LearningabstractFor a long time, the local descriptors learning benefited from the use of L2 normalization, which projects the descriptor space onto the hypersphere. However, there is no free lunch in the world. Although hypersphere description space stabilizes the optimization and improves the repeatability of the descriptors, it causes the descriptors to have a denser distribution, which reduces the discrimination between descriptors and leads to some incorrect matches. To alleviate this problem, we propose the learnablecross normalizationtechnology as an alternative to L2 normalization, which can achieve a consistent improvement in several of the current popular local descriptors. In addition, we propose an ER-Backbone that can efficiently reuse features in descriptors extraction and an IDC Loss that can provide an image-level description space distribution consistency constraint to further stimulate the performance of the local descriptors. Based on the above innovations, we provide a novel local descriptors extraction method named CNDesc. We perform experiments on image matching, homography estimation, 3D reconstruction, and visual localization tasks, and the results demonstrate that our CNDesc surpasses the current state-of-the-art local descriptors. Our code is available athttps://github.com/vignywang/CNDesc. Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | HTCViT: an effective network for image classification and segmentation based on natural disaster datasets
Wei Li 0237, Muyang Zhang, Weiliang Meng, Shibiao Xu, Xiaopeng Zhang 0001 |
Vis. Comput. | 6 |
| 2022 | MTLDesc: Looking Wider to Describe BetterabstractLimited by the locality of convolutional neural networks, most existing local features description methods only learn local descriptors with local information and lack awareness of global and surrounding spatial context. In this work, we focus on making local descriptors ``look wider to describe better'' by learning local Descriptors with More Than Local information (MTLDesc). Specifically, we resort to context augmentation and spatial attention mechanism to make the descriptors obtain non-local awareness. First, Adaptive Global Context Augmented Module and Diverse Local Context Augmented Module are proposed to construct robust local descriptors with context information from global to local. Second, we propose the Consistent Attention Weighted Triplet Loss to leverage spatial attention awareness in both optimization and matching of local descriptors. Third, Local Features Detection with Feature Pyramid is proposed to obtain more stable and accurate keypoints localization. With the above innovations, the performance of the proposed MTLDesc significantly surpasses the current state-of-the-art local descriptors on HPatches, Aachen Day-Night localization and InLoc indoor localization benchmarks. Our code is available at https://github.com/vignywang/MTLDesc. Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Bin Fan 0001, Xiaopeng Zhang 0001 |
AAAI | 7 |
| 2022 | DOMAINDESC: Learning Local Descriptors With Domain AdaptationabstractRobust and efficient local descriptor is crucial in a wide range of applications. In this paper, we propose a novel descriptor DomainDesc which is invariant as much as possible by learning local Descriptor with Domain adaptation. We design the feature-level domain adaptation loss to improve robustness of our DomainDesc by punishing inconsistent high-level feature distributions of different images, while we present the pixel-level cross-domain consistency loss to compensate for the inconsistency between the descriptors corresponding to the keypoints at the pixel level. Besides, we adopt a new architecture to make the descriptor contain as much information as possible, and combine triplet loss and cross-domain consistency loss for descriptor supervision to ensure the distinguished ability of our descriptor. Finally, we give a cross-domain dataset generation strategy to quickly construct our training dataset for diverse domains to adapt to complex application scenarios. Experiments validate that our DomainDesc achieves state-of-the-art performances on HPatches image matching benchmark and Aachen-Day-Night localization benchmark. Rongtao Xu, Changwei Wang 0001, Bin Fan 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001 |
ICASSP | 7 |
| 2022 | Softgan: Towards Accurate Lung Nodule Segmentation via Soft Mask SupervisionabstractAccurate lung nodule segmentation from Computed Tomog-raphy (CT) images is crucial to the analysis and diagnosis of lung diseases such as COVID-19 and lung cancer. How-ever, due to the variety of lung nodules and the lack of high-quality labeling, accurate lung nodule segmentation is still a challenging problem. In this paper, we propose a novel paradigm including an automatic accurate annotation pipeline and a segmentation network for this task. First, we introduce a new segmentation mask representation named Soft Mask which has richer and more accurate edge details description and better visualization, and we design a universal automatic Soft Mask annotation pipeline to deal with different datasets. Besides, we provide a new challenging lung nodules segmen-tation dataset with traditional binarized masks and our soft masks for further studies. Second, we propose an effective network called SoftGAN that includes an improved back-bone and an adversarial training framework with Soft Mask, in order to improve the performance of accurate lung nodules segmentation. Extensive experiments validate that our Soft-GAN outperforms the state-of-the-art methods for accurate lung nodule segmentation. [Datasetrelease] Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Jun Xiao 0005, Qimin Peng, Xiaopeng Zhang 0001 |
ICME | 7 |
| 2022 | GeoROS: Georeferenced Real-time Orthophoto Stitching with Unmanned Aerial VehicleabstractSimultaneous orthophoto stitching during the flight of Unmanned Aerial Vehicles (UAV) can greatly promote the practicability and instantaneity of diverse applications such as emergency disaster rescue, digital agriculture, and cadastral survey, which is of remarkable interest in aerial photogrammetry. However, the inaccurately estimated camera poses and the intuitive fusion strategy of existing methods lead to misalignment and distortion artifacts in orthophoto mosaics. To address these issues, we propose a Georeferenced Real-time Orthophoto Stitching method (GeoROS), which can achieve efficient and accurate camera pose estimation through exploiting geolocation information in monocular visual simultaneous localization and mapping (SLAM) and fuse transformed images via orthogonality-preserving criterion. Specifically, in the SLAM process, georeferenced tracking is employed to acquire high-quality initial camera poses with a geolocation based motion model and facilitate non-linear pose optimization. Meanwhile, we design a georeferenced mapping scheme by introducing robust geolocation constraints in joint optimization of camera poses and the position of landmarks. Finally, aerial images warped with localized cameras are fused by considering both the orthogonality of camera orientation relative to the ground plane and the pixel centrality to fulfill global orthorectification. Besides, we construct two datasets with global navigation satellite system (GNSS) information of different scenarios and validate the superiority of our GeoROS method compared with state-of-the-art methods in accuracy and efficiency. Guangze Gao, Mengke Yuan, Jiaming Gu, Weiliang Meng, Shibiao Xu, Xiaopeng Zhang 0001 |
IROS | 7 |
| 2022 | DA-Net: Dual Branch Transformer and Adaptive Strip Upsampling for Retinal Vessels Segmentation
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001 |
MICCAI (2) | 5 |
| 2022 | Scene text removal via cascaded text stroke detection and erasingabstractRecent learning-based approaches show promising performance improvement for the scene text removal task but usually leave several remnants of text and provide visually unpleasant results. In this work, a novel end-to-end framework is proposed based on accurate text stroke detection. Specifically, the text removal problem is decoupled into text stroke detection and stroke removal; we design separate networks to solve these two subproblems, the latter being a generative network. These two networks are combined as a processing unit, which is cascaded to obtain our final model for text removal. Experimental results demonstrate that the proposed method substantially outperforms the state-of-the-art for locating and erasing scene text. A new large-scale real-world dataset with 12,120 images has been constructed and is being made available to facilitate research, as current publicly available datasets are mainly synthetic so cannot properly measure the performance of different methods. Xuewei Bian, Weize Quan, Juntao Ye, Xiaopeng Zhang 0001, Dong-Ming Yan 0001 |
Comput. Vis. Media | 5 |
| 2022 | Instance segmentation of biological images using graph convolutional network
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001 |
Eng. Appl. Artif. Intell. | 6 |
| 2022 | Efficient Pairwise 3-D Registration of Urban Scenes via Hybrid Structural DescriptorsabstractAutomatic registration of point clouds captured by terrestrial laser scanning (TLS) plays an important role in many fields including remote sensing (e.g., transportation management, 3-D reconstruction in large-scale urban areas and environment monitoring), computer vision, and virtual reality and robotics. However, noise, outliers, nonuniform point density, and small overlaps are inevitable when collecting multiple views of data, which poses great challenges to 3-D registration of point clouds. Since conventional registration methods aim to find point correspondences and estimate transformation parameters directly in the original point space, the traditional way to address these difficulties is to introduce many restrictions during the scanning process (e.g., more scanning and careful selection of scanning positions), thus making the data acquisition more difficult. In this article, we present a novel 3-D registration framework that performs in a “middle-level structural space” and is capable of robustly and efficiently reconstructing urban, semiurban, and indoor scenes, despite disturbances introduced in the scanning process. The new structural space is constructed by extracting multiple types of middle-level geometric primitives (planes, spheres, cylinders, and cones) from the 3-D point cloud. We design a robust method to find effective primitive combinations corresponding to the 6-D poses of the raw point clouds and then construct hybrid-structure-based descriptors. By matching descriptors and computing rotation and translation parameters, successful registration is achieved. Note that the whole process of our method is performed in the structural space, which has the advantages of capturing geometric structures (the relationship between primitives) and semantic features (primitive types and parameters) in larger fields. Experiments show that our method achieves state-of-the-art performance in several point cloud registration benchmark datasets at different scales and even obtains good registration results for data without overlapping areas. Jianwei Guo 0003, Zhanglin Cheng, Jun Xiao 0005, Xiaopeng Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Progressive Feature Learning for Facade Parsing With OcclusionsabstractExisting deep models for facade parsing often fail in classifying pixels in heavily occluded regions of facade images due to the difficulty in feature representation of these pixels. In this paper, we solve facade parsing with occlusions by progressive feature learning. To this end, we locate the regions contaminated by occlusions via Bayesian uncertainty evaluation on categorizing each pixel in these regions. Then, guided by the uncertainty, we propose an occlusion-immune facade parsing architecture in which we progressively re-express the features of pixels in each contaminated region from easy to hard. Specifically, the outside pixels, which have reliable context from visible areas, are re-expressed at early stages; the inner pixels are processed at late stages when their surroundings have been decontaminated at the earlier stages. In addition, at each stage, instead of using regular square convolution kernels, we design a context enhancement module (CEM) with directional strip kernels, which can aggregate structural context to re-express facade pixels. Extensive experiments on popular facade datasets demonstrate that the proposed method achieves state-of-the-art performance. Wenguang Ma, Shibiao Xu, Wei Ma 0008, Xiaopeng Zhang 0001, Hongbin Zha |
IEEE Trans. Image Process. | 4 |
| 2022 | Single-Image Specular Highlight Removal via Real-World Dataset ConstructionabstractSpecular reflections pose great challenges on various multimedia and computer vision tasks,e.g., image segmentation, detection and matching. In this paper, we build a large-scale Paired Specular-Diffuse (PSD) image dataset, where the images are carefully captured by using real-world objects and the ground-truth specular-free diffuse images are provided. To the best of our knowledge, this is the first real-world benchmark dataset for specular highlight removal task, which is useful for evaluating and encouraging new deep learning-based approaches. Given this dataset, we present a novel Generative Adversarial Network (GAN) for specular highlight removal from a single image by introducing the detection of specular reflection information as a guidance. Our network also makes full use of the attention mechanism and is able to directly model the mapping relation between the diffuse area and the specular highlight area without any explicit estimation of the illumination. Experimental results demonstrate that the proposed network is more effective to remove specular reflection components with the guidance of specular highlight detection than recent state-of-the-art methods. Zhongqi Wu, Chuanqing Zhuang, Jianwei Guo 0003, Jun Xiao 0005, Xiaopeng Zhang 0001, Dong-Ming Yan 0001 |
IEEE Trans. Multim. | 6 |
| 2022 | Surface Remeshing: A Systematic Literature Review of Methods and Research DirectionsabstractTriangle meshes are used in many important shape-related applications including geometric modeling, animation production, system simulation, and visualization. However, these meshes are typically generated in raw form with several defects and poor-quality elements, obstructing them from practical application. Over the past decades, different surface remeshing techniques have been presented to improve these poor-quality meshes prior to the downstream utilization. A typical surface remeshing algorithm converts an input mesh into a higher quality mesh with consideration of given quality requirements as well as an acceptable approximation to the input mesh. In recent years, surface remeshing has gained significant attention from researchers and engineers, and several remeshing algorithms have been proposed. However, there has been no survey article on remeshing methods in general with a defined search strategy and article selection mechanism covering the recent approaches in surface remeshing domain with a good connection to classical approaches. In this article, we present a survey on surface remeshing techniques, classifying all collected articles in different categories and analyzing specific methods with their advantages, disadvantages, and possible future improvements. Following the systematic literature review methodology, we define step-by-step guidelines throughout the review process, including search strategy, literature inclusion/exclusion criteria, article quality assessment, and data extraction. With the aim of literature collection and classification based on data extraction, we summarized collected articles, considering the key remeshing objectives, the way the mesh quality is defined and improved, and the way their techniques are compared with other previous methods. Remeshing objectives are described by angle range control, feature preservation, error control, valence optimization, and remeshing compatibility. The metrics used in the literature for the evaluation of surface remeshing algorithms are discussed. Meshing techniques are compared with other related methods via a comprehensive table with indices of the method name, the remeshing challenge met and solved, the category the method belongs to, and the year of publication. We expect this survey to be a practical reference for surface remeshing in terms of literature classification, method analysis, and future prospects. Dawar Khan, Alexander Plopski, Yuichiro Fujimoto, Masayuki Kanbara, Gul Jabeen, Yongjie Jessica Zhang, Xiaopeng Zhang 0001, Hirokazu Kato 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2022 | Blending Surface Segmentation and Editing for 3D ModelsabstractRecognizing and fitting shape primitives from underlying 3D models are key components of many computer graphics and computer vision applications. Although a vast number of structural recovery methods are available, they usually fail to identify blending surfaces, which corresponds to small transitional regions among relatively large primary patches. To address this issue, we present a novel approach for automatic segmentation and surface fitting with accurate geometric parameters from 3D models, especially mechanical parts. Overall, we formulate the structural segmentation as a Markov random field (MRF) labeling problem. In contrast to existing techniques, we first propose a new clustering algorithm to build superfacets by incorporating 3D local geometric information. This algorithm extracts the general quadric and rolling-ball blending regions, and improves the robustness of further segmentation. Next, we apply a specially designed MRF framework to efficiently partition the original model into different meaningful patches of known surface types by defining the multilabel energy function on the superfacets. Furthermore, we present an iterative optimization algorithm based on skeleton extraction to fit rolling-ball blending patches by recovering the parameters of the rolling center trajectories and ball radius. Experiments on different complex models demonstrate the effectiveness and robustness of the proposed method, and the superiority of our method is also verified through comparisons with state-of-the-art approaches. We further apply our algorithm in applications such as mesh editing by changing the radius of the rolling balls. Jianwei Guo 0003, Jun Xiao 0005, Xiaopeng Zhang 0001, Dong-Ming Yan 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Triple-strip attention mechanism-based natural disaster images classification and segmentation
Mengke Yuan, Jiaming Gu, Weiliang Meng, Shibiao Xu, Xiaopeng Zhang 0001 |
Vis. Comput. | 6 |
| 2021 | Towards Effective Adversarial Attack Against 3D Point Cloud ClassificationabstractIn the domain of 3D point cloud classification, deep learning based classifiers have made significant progress, while they have been also proven to be vulnerable on the adversarial at-tack at the same time. Some recent works employ the attack methods that devised for image classification such as projected gradient descent (PGD) to attack the 3D classifiers, but their performances seem quite limited when faced with statistical operations including point cloud denoising and point cloud upsampling. In this paper, we propose ‘SmoothAttack’, a new attack that can craft adversarial point clouds robust to statistical operations. SmoothAttack can be easily applied in both global constraint and pointwise constraint. Besides, we analyze the directions of perturbations onto the point cloud during the iteration process, where SmoothAttack can some-how stabilize the direction and make full use of the adversarial budgets. Experiments validate that our ‘SmoothAttack’ can raise the attack success rates against statistical defenses up to 98% for untargeted attack and 91% for targeted attack on ModelNet40 database when fooling the classifiers Point-Net and DGCNN. Chengcheng Ma, Weiliang Meng, Baoyuan Wu, Shibiao Xu, Xiaopeng Zhang 0001 |
ICME | 5 |
| 2021 | DC-Net: Dual Context Network for 2D Medical Image Segmentation
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001 |
MICCAI (1) | 5 |
| 2021 | Extracting Cycle-aware Feature Curve Networks from 3D Models
Zhengda Lu, Jianwei Guo 0003, Jun Xiao 0005, Ying Wang 0030, Xiaopeng Zhang 0001, Dong-Ming Yan 0001 |
Comput. Aided Des. | 5 |
| 2021 | Customized Summarizations of Visual Data CollectionsabstractAbstract We propose a framework to generate customized summarizations of visual data collections, such as collections of images, materials, 3D shapes, and 3D scenes. We assume that the elements in the visual data collections can be mapped to a set of vectors in a feature space, in which a fitness score for each element can be defined, and we pose the problem of customized summarizations as selecting a subset of these elements. We first describe the design choices a user should be able to specify for modeling customized summarizations and propose a corresponding user interface. We then formulate the problem as a constrained optimization problem with binary variables and propose a practical and fast algorithm based on the alternating direction method of multipliers (ADMM). Our results show that our problem formulation enables a wide variety of customized summarizations, and that our solver is both significantly faster than state‐of‐the‐art commercial integer programming solvers and produces better solutions than fast relaxation‐based solvers. Mengke Yuan, Bernard Ghanem, Dong-Ming Yan 0001, Baoyuan Wu, Xiaopeng Zhang 0001, Peter Wonka |
Comput. Graph. Forum | 5 |
| 2021 | Data-driven floor plan understanding in rural residential buildings via deep recognition
Zhengda Lu, Jianwei Guo 0003, Weiliang Meng, Jun Xiao 0005, Xiaopeng Zhang 0001 |
Inf. Sci. | 7 |
| 2021 | Multi-Feature Super-Resolution Network for Cloth Wrinkle Synthesis
Juntao Ye, Xiaopeng Zhang 0001 |
J. Comput. Sci. Technol. | 3 |
| 2021 | Efficient Center Voting for Object Detection and 6D Pose Estimation in 3D Point CloudabstractWe present a novel and efficient approach to estimate 6D object poses of known objects in complex scenes represented by point clouds. Our approach is based on the well-known point pair feature (PPF) matching, which utilizes self-similar point pairs to compute potential matches and thereby cast votes for the object pose by a voting scheme. The main contribution of this paper is to present an improved PPF-based recognition framework, especially a new center voting strategy based on the relative geometric relationship between the object center and point pair features. Using this geometric relationship, we first generate votes to object centers resulting in vote clusters near real object centers. Then we group and aggregate these votes to generate a set of pose hypotheses. Finally, a pose verification operator is performed to filter out false positives and predict appropriate 6D poses of the target object. Our approach is also suitable to solve the multi-instance and multi-object detection tasks. Extensive experiments on a variety of challenging benchmark datasets demonstrate that the proposed algorithm is discriminative and robust towards similar-looking distractors, sensor noise, and geometrically simple shapes. The advantage of our work is further verified by comparing to the state-of-the-art approaches. Jianwei Guo 0003, Xuejun Xing, Weize Quan, Dong-Ming Yan 0001, Qingyi Gu, Yang Liu 0014, Xiaopeng Zhang 0001 |
IEEE Trans. Image Process. | 7 |
| 2021 | TreePartNet: neural decomposition of point clouds for 3D tree reconstructionabstractWe present TreePartNet , a neural network aimed at reconstructing tree geometry from point clouds obtained by scanning real trees. Our key idea is to learn a natural neural decomposition exploiting the assumption that a tree comprises locally cylindrical shapes. In particular, reconstruction is a two-step process. First, two networks are used to detect priors from the point clouds. One detects semantic branching points, and the other network is trained to learn a cylindrical representation of the branches. In the second step, we apply a neural merging module to reduce the cylindrical representation to a final set of generalized cylinders combined by branches. We demonstrate results of reconstructing realistic tree geometry for a variety of input models and with varying input point quality, e.g., noise, outliers, and incompleteness. We evaluate our approach extensively by using data from both synthetic and real trees and comparing it with alternative methods. Jianwei Guo 0003, Bedrich Benes, Oliver Deussen, Xiaopeng Zhang 0001, Hui Huang 0004 |
ACM Trans. Graph. | 5 |
| 2020 | MLIFeat: Multi-level Information Fusion Based Deep Local Features
Jinge Wang 0006, Shibiao Xu, Xiaopeng Zhang 0001 |
ACCV (3) | 5 |
| 2020 | Efficient Joint Gradient Based Attack Against SOR Defense for 3D Point Cloud ClassificationabstractDeep learning based classifiers on 3D point cloud data have been shown vulnerable to adversarial examples, while a defense strategy named Statistical Outlier Removal (SOR) is widely adopted to defend adversarial examples successfully, by discarding outlier points in the point cloud. Chengcheng Ma, Weiliang Meng, Baoyuan Wu, Shibiao Xu, Xiaopeng Zhang 0001 |
ACM Multimedia | 5 |
| 2020 | Learning local shape descriptors for computing non-rigid dense correspondenceabstractA discriminative local shape descriptor plays an important role in various applications. In this paper, we present a novel deep learning framework that derives discriminative local descriptors for deformable 3D shapes. We use local “geometry images” to encode the multi-scale local features of a point, via an intrinsic parameterization method based on geodesic polar coordinates. This new parameterization provides robust geometry images even for badly-shaped triangular meshes. Then a triplet network with shared architecture and parameters is used to perform deep metric learning; its aim is to distinguish between similar and dissimilar pairs of points. Additionally, a newly designed triplet loss function is minimized for improved, accurate training of the triplet network. To solve the dense correspondence problem, an efficient sampling approach is utilized to achieve a good compromise between training performance and descriptor quality. During testing, given a geometry image of a point of interest, our network outputs a discriminative local descriptor for it. Extensive testing of non-rigid dense shape matching on a variety of benchmarks demonstrates the superiority of the proposed descriptors over the state-of-the-art alternatives. Jianwei Guo 0003, Hanyu Wang 0002, Zhanglin Cheng, Xiaopeng Zhang 0001, Dong-Ming Yan 0001 |
Comput. Vis. Media | 4 |
| 2020 | Learning across views for stereo image completionabstractStereo image completion (SIC) is to fill holes existing in a pair of stereo images. SIC is more complicated than single image repairing, which needs to complete the pair of images while keeping their stereoscopic consistency. In recent years, deep learning has been introduced into single image repairing but seldom used for SIC. The authors present a novel deep learning‐based approach for SIC. In their method, an X‐shaped fully convolutional network (called SICNet) is proposed and designed to complete stereo images, which is composed of two branches of convolutional neural network layers to encode the context of the left and right images separately, a fusion module for stereo‐interactive completion, and two branches of decoders to produce completed left and right images, respectively. In consideration of both inter‐view and intra‐view cues, they introduce auxiliary networks and define comprehensive losses to train SICNet to perform single‐view coherent and cross‐view consistent completion simultaneously. Extensive experiments are conducted to show the state‐of‐the‐art performances of the proposed approach and its key components. Wei Ma 0008, Mana Zheng, Wenguang Ma, Shibiao Xu, Xiaopeng Zhang 0001 |
IET Comput. Vis. | 5 |
| 2020 | High accuracy correspondence field estimation via MST based patch matching
Feihu Zhang, Shibiao Xu, Xiaopeng Zhang 0001 |
Multim. Tools Appl. | 3 |
| 2020 | Unsupervised Multi-View Constrained Convolutional Network for Accurate Depth EstimationabstractAccurate depth estimation from images is a fundamental problem in computer vision. In this paper, we propose an unsupervised learning based method to predict high-quality depth map from multiple images. A novel multi-view constrained DenseDepthNet is designed for this task. Our DenseDepthNet can effectively leverage both the low-level and high-level features of input images and generate appealing results, especially with sharp details. We employ the public datasets KITTI and Cityscapes for training in an end-to-end unsupervised fashion. A novel depth consistency loss based on multi-view geometry constraint is also applied to the corresponding points across pairwise images, which helps to improve the quality of predicted depth maps significantly. We conduct comprehensive evaluations on our DenseDepthNet and our depth consistency loss function. Experiments validate that our method outperforms the state-of-the-art unsupervised methods and produce comparable results with supervised methods. Shibiao Xu, Baoyuan Wu, Weiliang Meng, Xiaopeng Zhang 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | Inverse Procedural Modeling of Branching Structures by Inferring L-SystemsabstractWe introduce an inverse procedural modeling approach that learns L-system representations of pixel images with branching structures. Our fully automatic model generates a compact set of textual rewriting rules that describe the input. We use deep learning to discover atomic structures such as line segments or branchings. Orientation and scaling of these structures are determined and the detected structures are combined into a tree. The initial representation is analyzed, and repeating parts are encoded into a small grammar by using greedy optimization while the user can control the size of the detected rules. The output is an L-system that represents the input image as a simple text and a set of terminal symbols. We apply our approach to a variety of examples, demonstrate its robustness against noise and blur, and we show that it can detect user sketches and complex input structures. Jianwei Guo 0003, Haiyong Jiang, Bedrich Benes, Oliver Deussen, Xiaopeng Zhang 0001, Dani Lischinski, Hui Huang 0004 |
ACM Trans. Graph. | 5 |
| 2020 | MGCN: descriptor learning using multiscale GCNsabstractWe propose a novel framework for computing descriptors for characterizing points on three-dimensional surfaces. First, we present a new non-learned feature that uses graph wavelets to decompose the Dirichlet energy on a surface. We call this new feature Wavelet Energy Decomposition Signature (WEDS). Second, we propose a new Multiscale Graph Convolutional Network (MGCN) to transform a non-learned feature to a more discriminative descriptor. Our results show that the new descriptor WEDS is more discriminative than the current state-of-the-art non-learned descriptors and that the combination of WEDS and MGCN is better than the state-of-the-art learned descriptors. An important design criterion for our descriptor is the robustness to different surface discretizations including triangulations with varying numbers of vertices. Our results demonstrate that previous graph convolutional networks significantly overfit to a particular resolution or even a particular triangulation, but MGCN generalizes well to different surface discretizations. In addition, MGCN is compatible with previous descriptors and it can also be used to improve the performance of other descriptors, such as the heat kernel signature, the wave kernel signature, or the local point signature. Yiqun Wang 0001, Jing Ren 0004, Dong-Ming Yan 0001, Jianwei Guo 0003, Xiaopeng Zhang 0001, Peter Wonka |
ACM Trans. Graph. | 5 |
| 2020 | Realistic Procedural Plant Modeling from Multiple View ImagesabstractIn this paper, we describe a novel procedural modeling technique for generating realistic plant models from multi-view photographs. The realism is enhanced via visual and spatial information acquired from images. In contrast to previous approaches that heavily rely on user interaction to segment plants or recover branches in images, our method automatically estimates an accurate depth map of each image and extracts a 3D dense point cloud by exploiting an efficient stereophotogrammetry approach. Taking this point cloud as a soft constraint, we fit a parametric plant representation to simulate the plant growth progress. In this way, we are able to synthesize parametric plant models from real data provided by photos and 3D point clouds. We demonstrate the robustness of the proposed approach by modeling various plants with complex branching structures and significant self-occlusions. We also demonstrate that the proposed framework can be used to reconstruct ground-covering plants, such as bushes and shrubs which have been given little attention in the literature. The effectiveness of our approach is validated by visually and quantitatively comparing with the state-of-the-art approaches. Jianwei Guo 0003, Shibiao Xu, Dong-Ming Yan 0001, Zhanglin Cheng, Marc Jaeger 0002, Xiaopeng Zhang 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2020 | Selection Expressions for Procedural ModelingabstractWe introduce a new approach for procedural modeling. Our main idea is to select shapes using selection-expressions instead of simple string matching used in current state-of-the-art grammars like CGA shape and CGA++. A selection-expression specifies how to select a potentially complex subset of shapes from a shape hierarchy, e.g., "select all tall windows in the second floor of the main building facade". This new way of modeling enables us to express modeling ideas in their global context rather than traditional rules that operate only locally. To facilitate selection-based procedural modeling we introduce the procedural modeling language SelEx. An important implication of our work is that enforcing important constraints, such as alignment and same size constraints can be done by construction. Therefore, our procedural descriptions can generate facade and building variations without violating alignment and sizing constraints that plague the current state of the art. While the procedural modeling of architecture is our main application domain, we also demonstrate that our approach nicely extends to other man-made objects. Haiyong Jiang, Dong-Ming Yan 0001, Xiaopeng Zhang 0001, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2019 | A Robust Local Spectral Descriptor for Matching Non-Rigid Shapes With Incompatible Shape StructuresabstractConstructing a robust and discriminative local descriptor for 3D shape is a key component of many computer vision applications. Although existing learning-based approaches can achieve good performance in some specific benchmarks, they usually fail to learn enough information from shapes with different shape types and structures (e.g., spatial resolution, connectivity, transformations, etc.) Focusing on this issue, in this paper, we present a more discriminative local descriptor for deformable 3D shapes with incompatible structures. Based on the spectral embedding using the Laplace-Beltrami framework on the surface, we first construct a novel local spectral feature which shows great resilience to change in mesh resolution, triangulation, transformation. Then the multi-scale local spectral features around each vertex are encoded into a `geometry image', called vertex spectral image, in a very compact way. Such vertex spectral images can be efficiently trained to learn local descriptors using a triplet neural network. Finally, for training and evaluation, we present a new benchmark dataset by extending the widely used FAUST dataset. We utilize a remeshing approach to generate modified shapes with different structures. We evaluate the proposed approach thoroughly and make an extensive comparison to demonstrate that our approach outperforms recent state-of-the-art methods on this benchmark. Yiqun Wang 0001, Jianwei Guo 0003, Dong-Ming Yan 0001, Kai Wang 0002, Xiaopeng Zhang 0001 |
CVPR | 5 |
| 2019 | Fast and Error-Bounded Space-Variant Bilateral Filtering
Mengke Yuan, Longquan Dai, Dong-Ming Yan 0001, Liqiang Zhang 0001, Jun Xiao 0005, Xiaopeng Zhang 0001 |
J. Comput. Sci. Technol. | 6 |
| 2019 | Parameter optimization criteria guided 3D point cloud classification
Hongjun Li 0002, Weiliang Meng, Shiming Xiang, Xiaopeng Zhang 0001 |
Multim. Tools Appl. | 5 |
| 2019 | Joint face alignment and segmentation via deep multi-task learning
Fan Tang, Weiming Dong, Feiyue Huang, Xiaopeng Zhang 0001 |
Multim. Tools Appl. | 5 |
| 2019 | Interpreting and Extending the Guided Filter via Cyclic Coordinate DescentabstractThe guided filter (GF) is a widely used smoothing tool in computer vision and image processing. However, to the best of our knowledge, few papers investigate the mathematical connection between this filter and the least-squares optimization. In this paper, we first interpret the guided filter as the cyclic coordinate descent (CCD) solver of a least-squares objective function. This discovery implies an extension approach to generalize the guided filter since we can change the least-squares objective function and define new filters as the first pass iteration of the CCD solver of modified objective functions. In addition, referring to the iterative minimizing procedure of the CCD, we can derive new rolling filtering schemes. So, we are reasonable to say that our discovery not only reveals an approach to design new GF-like filters adapting to specific requirements of applications but also offers thorough explanations for two rolling filtering schemes of the guided filter as well as the method to extend them. Experiments prove our new proposed filters and rolling filtering schemes could produce state-of-the-art results. Longquan Dai, Mengke Yuan, Yuan Xie 0006, Xiaopeng Zhang 0001, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 5 |
| 2019 | Isotropic Surface Remeshing without Large and Small AnglesabstractWe introduce a novel algorithm for isotropic surface remeshing which progressively eliminates obtuse triangles and improves small angles. The main novelty of the proposed approach is a simple vertex insertion scheme that facilitates the removal of large angles, and a vertex removal operation that improves the distribution of small angles. In combination with other standard local mesh operators, e.g., connectivity optimization and local tangential smoothing, our algorithm is able to remesh efficiently a low-quality mesh surface. Our approach can be applied directly or used as a post-processing step following other remeshing approaches. Our method has a similar computational efficiency to the fastest approach available, i.e., real-time adaptive remeshing [1]. In comparison with state-of-the-art approaches, our method consistently generates better results based on evaluations using different metrics. Yiqun Wang 0001, Dong-Ming Yan 0001, Chengcheng Tang, Jianwei Guo 0003, Xiaopeng Zhang 0001, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2018 | Large-scale 3D Point Cloud Classification Based On Feature Description Matrix By CNNabstractLarge-scale 3D Point cloud classification is a basic topic for various applications. Traditional geometries features are usually independent of each other and difficult to adapt to a fixed classification model. With the rise of the neural network, deep learning is considered in 3D point cloud application. 3D points are difficult to feed the neural network directly based on deep learning, as they cannot be arranged in a fixed order as image pixels. In this paper, we combine traditional feature-based methods with the Convolutional neural network(CNN) to finish the classification task. The core idea is to construct a feasible structure called Feature Description Matrix(FDM) which encapsulates the local feature of the point to feed CNN for training and testing. By extracting geometry features and designed Feature Description Vectors(FDV) for FDM, a simple mechanism for point cloud classification is given, and experiments validate the effectiveness of our method, with higher classification accuracy compared to state-of-art works. Lei Wang 0089, Weiliang Meng, Runping Xi, Yanning Zhang 0001, Ling Lu, Xiaopeng Zhang 0001 |
CASA | 6 |
| 2018 | Learning 3D Keypoint Descriptors for Non-rigid Shape Matching
Hanyu Wang 0002, Jianwei Guo 0003, Dong-Ming Yan 0001, Weize Quan, Xiaopeng Zhang 0001 |
ECCV (8) | 5 |
| 2018 | Photo Squarization by Deep Multi-Operator RetargetingabstractSquared forms of photos are widely used in social media as album covers or thumbnails of image streams. In this study, we realize photo squarization by modeling Retargeting Visual Perception Issues, which reflect human perception preference toward image ratargeting. General image retargeting techniques deal with three common issues, namely, salient content, object shape, and scene composition, to preserve the important information of original image. We propose a new way based on multi-operator techniques to investigate human behavior in balancing the three issues. We establish a new dataset and observe human behavior by inviting investigators to retarget images to square manually. We propose a data-driven approach composed of perception and distillation modules by using deep learning techniques to predict human perception preference. The perception part learns the relations among the three issues, and the distillation part transfers the learned relations to a simple but effective network. Our study contributes to deep learning literature by optimizing a network index and lightening its running burden. Experimental results show that photo squarization results generated by the proposed model are consistent with human visual perception results. Fan Tang, Weiming Dong, Xiaopeng Zhang 0001, Oliver Deussen, Tong-Yee Lee |
ACM Multimedia | 4 |
| 2018 | Learning Scene Illumination by Pairwise Photos from Rear and Front Mobile CamerasabstractAbstract Illumination estimation is an essential problem in computer vision, graphics and augmented reality. In this paper, we propose a learning based method to recover low‐frequency scene illumination represented as spherical harmonic (SH) functions by pairwise photos from rear and front cameras on mobile devices. An end‐to‐end deep convolutional neural network (CNN) structure is designed to process images on symmetric views and predict SH coefficients. We introduce a novel Render Loss to improve the rendering quality of the predicted illumination. A high quality high dynamic range (HDR) panoramic image dataset was developed for training and evaluation. Experiments show that our model produces visually and quantitatively superior results compared to the state‐of‐the‐arts. Moreover, our method is practical for mobile‐based applications. Dachuan Cheng, Yanyun Chen, Xiaoming Deng 0001, Xiaopeng Zhang 0001 |
Comput. Graph. Forum | 5 |
| 2018 | Tree Growth Modelling Constrained by Growth EquationsabstractAbstract Modelling and simulation of tree growth that is faithful to the living environment and numerically consistent to botanic knowledge are important topics for realistic modelling in computer graphics. The realism factors concerned include the effects of complex environment on tree growth and the reliability of the simulation in botanical research, such as horticulture and agriculture. This paper proposes a new approach, namely, integrated growth modelling, to model virtual trees and simulate their growth by enforcing constraints of environmental resources and tree morphological properties. Morphological properties are integrated into a growth equation with different parameters specified in the simulation, including its sensitivity to light, allocation and usage of received resources and effects on its environment. The growth equation guarantees that the simulation procedure numerically matches the natural growth phenomenon of trees. With this technique, the growth procedures of diverse and realistic trees can also be modelled in different environments, such as resource competition among multiple trees. Lei Yi, Hongjun Li 0002, Jianwei Guo 0003, Oliver Deussen, Xiaopeng Zhang 0001 |
Comput. Graph. Forum | 5 |
| 2018 | Surface remeshing with robust user-guided segmentationabstractSurface remeshing is widely required in modeling, animation, simulation, and many other computer graphics applications. Improving the elements’ quality is a challenging task in surface remeshing. Existing methods often fail to efficiently remove poor-quality elements especially in regions with sharp features. In this paper, we propose and use a robust segmentation method followed by remeshing the segmented mesh. Mesh segmentation is initiated using an existing Live-wire interaction approach and is further refined using local mesh operations. The refined segmented mesh is finally sent to the remeshing pipeline, in which each mesh segment is remeshed independently. An experimental study compares our mesh segmentation method as well as remeshing results with representative existing methods. We demonstrate that the proposed segmentation method is robust and suitable for remeshing. Dawar Khan, Dong-Ming Yan 0001, Yixin Zhuang, Xiaopeng Zhang 0001 |
Comput. Vis. Media | 5 |
| 2018 | Synthesizing cloth wrinkles by CNN-based geometry image superresolutionabstractAbstract We propose a novel deep learning‐based method, called mesh superresolution, to enrich low‐resolution (LR) cloth meshes with wrinkles. A pair of low and high‐resolution (HR) meshes are simulated, with the simulation of the HR mesh tracks with that of the LR mesh. The frame data are converted into geometry images and used as a training data set. A residual network, called SR residual network, is employed to train an image synthesizer that superresolves an LR image into an HR one. Once the HR image is converted back to an HR mesh, it is abundant in wrinkles compared with its coarse counterpart. The synthesizing is very efficient and is 24× faster than a full HR simulation. We demonstrate the performances of mesh superresolution with various simulation scenes. Juntao Ye, Liguo Jiang, Chengcheng Ma, Zhanglin Cheng, Xiaopeng Zhang 0001 |
Comput. Animat. Virtual Worlds | 6 |
| 2018 | Interactive stereo image segmentation via adaptive prior selection
Wei Ma 0008, Shibiao Xu, Xiaopeng Zhang 0001 |
Multim. Tools Appl. | 4 |
| 2018 | Accurate blind deblurring using salientpatch-based prior for large-size images
Chengcheng Ma, Jiguang Zhang, Shibiao Xu, Weiliang Meng, Runping Xi, G. Hemanth Kumar, Xiaopeng Zhang 0001 |
Multim. Tools Appl. | 7 |
| 2018 | Real-time pedestrian detection via hierarchical convolutional feature
Dongming Yang, Jiguang Zhang, Shibiao Xu, Shuiying Ge, G. Hemanth Kumar, Xiaopeng Zhang 0001 |
Multim. Tools Appl. | 6 |
| 2018 | Automatic Building Rooftop Extraction From Aerial Images via Hierarchical RGB-D PriorsabstractAccurate building rooftop extraction from high-resolution aerial images is of crucial importance in a wide range of applications. Owing to the varying appearance and large-scale range of scene objects, especially for building rooftops in different scales and heights, single-scale or individual prior-based extraction technique is insufficient in pursuing efficient, generic, and accurate extraction results. The trend toward integrating multiscale or several cue techniques appears to be the best way; thus, such integration is the focus of this paper. We first propose a novel salient rooftop detector integrating four correlative RGB-D priors (depth cue, uniqueness prior, shape prior, and transition surface prior) for improved rooftop extraction to address the preceding complex issues mentioned. Then, these correlative cues are computed from image layers created by our multilevel segmentation and further fused into the state-of-the-art high-order conditional random field (CRF) framework to locate the rooftop. Finally, an iterative optimization strategy is applied for high-quality solving, which can robustly handle varying appearance of building rooftops. Performance evaluations in the SZTAKI-INRIA benchmark data sets show that our method outperforms the traditional color-based algorithm and the original high-order CRF algorithm and its variants. The proposed algorithm is also evaluated and found to produce consistently satisfactory results for various large-scale, real-world data sets. Shibiao Xu, Xingjia Pan, Er Li, Baoyuan Wu, Shuhui Bu, Weiming Dong, Shiming Xiang, Xiaopeng Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2018 | Distinguishing Between Natural and Computer-Generated Images Using Convolutional Neural NetworksabstractDistinguishing between natural images (NIs) and computer-generated (CG) images by naked human eyes is difficult. In this paper, we propose an effective method based on a convolutional neural network (CNN) for this fundamental image forensic problem. Having observed the rather limited performance of training existing CCNs from scratch or fine-tuning pre-trained network, we design and implement a new and appropriate network with two cascaded convolutional layers at the bottom of a CNN. Our network can be easily adjusted to accommodate different sizes of input image patches while maintaining a fixed depth, a stable structure of CNN, and a good forensic performance. Considering the complexity of training CNNs and the specific requirement of image forensics, we introduce the so-called local-to-global strategy in our proposed network. Our CNN derives a forensic decision on local patches, and a global decision on a full-sized image can be easily obtained via simple majority voting. This strategy can also be used to improve the performance of existing methods that are based on hand-crafted features. Experimental results show that our method outperforms existing methods, especially in a challenging forensic scenario with NIs and CG images of heterogeneous origins. Our method also has good robustness against typical post-processing operations, such as resizing and JPEG compression. Unlike previous attempts to use CNNs for image forensics, we try to understand what our CNN has learned about the differences between NIs and CG images with the aid of adequate and advanced visualization tools. Weize Quan, Kai Wang 0002, Dong-Ming Yan 0001, Xiaopeng Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2018 | Animated Construction of Chinese Brush PaintingsabstractIn this paper, we present a method for reconstructing the drawing process of Chinese brush paintings. We demonstrate the possibility of computing an artistically reasonable drawing order from a static brush painting that is consistent with the rules of art. We map the key principles of drawing composition to our computational framework, which first organizes the strokes in three stages and then optimizes stroke ordering with natural evolution strategies. Our system produces reasonable animated constructions of Chinese brush paintings with minimal or no user intervention. We test our algorithm on a range of input paintings with varying degrees of complexity and structure and then evaluate the results via a user study. We discuss the applications of the proposed system to painting instruction, painting animation, and image stylization, especially in the context of art teaching. Fan Tang, Weiming Dong, Yiping Meng, Xing Mei, Feiyue Huang, Xiaopeng Zhang 0001, Oliver Deussen |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2017 | Hardware-Efficient Guided Image Filtering for Multi-label ProblemabstractThe Guided Filter (GF) is well-known for its linear complexity. However, when filtering an image with an n-channel guidance, GF needs to invert an n × n matrix for each pixel. To the best of our knowledge existing matrix inverse algorithms are inefficient on current hardwares. This shortcoming limits applications of multichannel guidance in computation intensive system such as multi-label system. We need a new GF-like filter that can perform fast multichannel image guided filtering. Since the optimal linear complexity of GF cannot be minimized further, the only way thus is to bring all potentialities of current parallel computing hardwares into full play. In this paper we propose a hardware-efficient Guided Filter (HGF), which solves the efficiency problem of multichannel guided image filtering and yields competent results when applying it to multi-label problems with synthesized polynomial multichannel guidance. Specifically, in order to boost the filtering performance, HGF takes a new matrix inverse algorithm which only involves two hardware-efficient operations: element-wise arithmetic calculations and box filtering. In order to break the linear model restriction, HGF synthesizes a polynomial multichannel guidance to introduce nonlinearity. Benefiting from our polynomial guidance and hardware-efficient matrix inverse algorithm, HGF not only is more sensitive to the underlying structure of guidance but also achieves the fastest computing speed. Due to these merits, HGF obtains state-of-the-art results in terms of accuracy and efficiency in the computation intensive multi-label systems. Longquan Dai, Mengke Yuan, Zechao Li, Xiaopeng Zhang 0001, Jinhui Tang 0001 |
CVPR | 4 |
| 2017 | A Unified Cloth Untangling Framework Through Discrete Collision DetectionabstractAbstract We present an efficient and stable framework, called Unified Intersection Resolver (UIR), for cloth simulation systems where not only impending collisions but also pre‐existing penetrations often arise. These two types of collisions are handled in a unified manner, by detecting edge‐face intersections first and then forming penetration stencils to be resolved iteratively. A stencil is a quadruple of vertices and it reveals either a vertex‐face or an edge‐edge collision event happened. Each quadruple also implicitly defines a collision normal, through which the four stencil vertices can be relocated, so that the corresponding edge‐face intersection disappear. We deduce three different ways, i.e., from predefined surface orientation, from history data and from global intersection analysis, to determine the collision normals of these stencils robustly. Multiple stencils that constitute a penetration region are processed simultaneously to eliminate penetrations. Cloth trapped in pinched environmental objects can be handled easily within our framework. We highlight its robustness by a number of challenging experiments involving collisions. Juntao Ye, Guanghui Ma, Liguo Jiang, Jituo Li, Gang Xiong 0001, Xiaopeng Zhang 0001 |
Comput. Graph. Forum | 7 |
| 2017 | Tree Branch Level of Detail Models for Forest NavigationabstractAbstract We present a level of detail (LOD) method designed for tree branches. It can be combined with methods for processing tree foliage to facilitate navigation through large virtual forests. Starting from a skeletal representation of a tree, we fit polygon meshes of various densities to the skeleton while the mesh density is adjusted according to the required visual fidelity. For distant models, these branch meshes are gradually replaced with semi‐transparent lines until the tree recedes to a few lines. Construction of these complete LOD models is guided by error metrics to ensure smooth transitions between adjacent LOD models. We then present an instancing technique for discrete LOD branch models, consisting of polygon meshes plus semi‐transparent lines. Line models with different transparencies are instanced on the GPU by merging multiple tree samples into a single model. Our technique reduces the number of draw calls in GPU and increases rendering performance. Our experiments demonstrate that large‐scale forest scenes can be rendered with excellent detail and shadows in real time. Xiaopeng Zhang 0001, Guanbo Bao, Weiliang Meng, Marc Jaeger 0002, Hongjun Li 0002, Oliver Deussen, Baoquan Chen |
Comput. Graph. Forum | 1 |
| 2017 | Orientation judgment for abstract paintings
Weiming Dong, Xiaopeng Zhang 0001, Zhiguo Jiang 0001 |
Multim. Tools Appl. | 3 |
| 2017 | Shape exploration of 3D heterogeneous models based on cages
Weiliang Meng, Jianwei Guo 0003, Xavier Bonaventura, Mateu Sbert, Xiaopeng Zhang 0001 |
Multim. Tools Appl. | 5 |
| 2017 | A Simple Push-Pull Algorithm for Blue-Noise SamplingabstractWe describe a simple push-pull optimization (PPO) algorithm for blue-noise sampling by enforcing spatial constraints on given point sets. Constraints can be a minimum distance between samples, a maximum distance between an arbitrary point and the nearest sample, and a maximum deviation of a sample's capacity (area of Voronoi cell) from the mean capacity. All of these constraints are based on the topology emerging from Delaunay triangulation, and they can be combined for improved sampling quality and efficiency. In addition, our algorithm offers flexibility for trading-off between different targets, such as noise and aliasing. We present several applications of the proposed algorithm, including anti-aliasing, stippling, and non-obtuse remeshing. Our experimental results illustrate the efficiency and the robustness of the proposed approach. Moreover, we demonstrate that our remeshing quality is superior to the current state-of-the-art approaches. Abdalla G. M. Ahmed, Jianwei Guo 0003, Dong-Ming Yan 0001, Jean-Yves Franceschi, Xiaopeng Zhang 0001, Oliver Deussen |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2016 | Tetrahedral meshing via maximal Poisson-disk sampling
Jianwei Guo 0003, Dong-Ming Yan 0001, Li Chen 0031, Xiaopeng Zhang 0001, Oliver Deussen, Peter Wonka |
Comput. Aided Geom. Des. | 4 |
| 2016 | Anisotropic Strain Limiting for Quadrilateral and Triangular Cloth MeshesabstractAbstract The cloth simulation systems often suffer from excessive extension on the polygonal mesh, so an additional strain‐limiting process is typically used as a remedy in the simulation pipeline. A cloth model can be discretized as either a quadrilateral mesh or a triangular mesh, and their strains are measured differently. The edge‐based strain‐limiting method for a quadrilateral mesh creates anisotropic behaviour by nature, as discretization usually aligns the edges along the warp and weft directions. We improve this anisotropic technique by replacing the traditionally used equality constraints with inequality ones in the mathematical optimization, and achieve faster convergence. For a triangular mesh, the state‐of‐the‐art technique measures and constrains the strains along the two principal (and constantly changing) directions in a triangle, resulting in an isotropic behaviour which prohibits shearing. Based on the framework of inequality‐constrained optimization, we propose a warp and weft strain‐limiting formulation. This anisotropic model is more appropriate for textile materials that do not exhibit isotropic strain behaviour. Guanghui Ma, Juntao Ye, Jituo Li, Xiaopeng Zhang 0001 |
Comput. Graph. Forum | 4 |
| 2016 | Symmetrization of facade layouts
Haiyong Jiang, Dong-Ming Yan 0001, Weiming Dong, Fuzhang Wu, Liangliang Nan, Xiaopeng Zhang 0001 |
Graph. Model. | 6 |
| 2016 | Analyzing surface sampling patterns using the localized pair correlation functionabstractPoint distributions with different characteristics have a crucial influence on graphics applications. Various analysis tools have been developed in recent years, mainly for blue noise sampling in Euclidean domains. In this paper, we present a new method to analyze the properties of general sampling patterns that are distributed on mesh surfaces. The core idea is to generalize to surfaces the pair correlation function (PCF) which has successfully been employed in sampling pattern analysis and synthesis in 2D and 3D. Experimental results demonstrate that the proposed approach can reveal correlations of point sets generated by a wide range of sampling algorithms. An acceleration technique is also suggested to improve the performance of the PCF. Weize Quan, Jianwei Guo 0003, Dong-Ming Yan 0001, Weiliang Meng, Xiaopeng Zhang 0001 |
Comput. Vis. Media | 5 |
| 2016 | Interactive Stereo Image Segmentation With RGB-D Hybrid ConstraintsabstractThis letter presents an approach to extracting a target object interactively from a given pair of stereo images. First, a user marks a few parts of the object and background in either of the two views with strokes. The marked pixels are used to generate the prior models of the foreground and background. Second, a graph is constructed with constraints formulated by the priors of foreground/background, similarities between intraview neighbor pixels and correspondences between interview pixels. Third, two segments of the foreground are extracted from the two views by optimization of the graph via graph cut. Traditional methods generally define the priors and neighbor similarities in RGB space. Differently, the proposed method integrates disparity distributions of foreground/background to enrich the priors and defines the similarity metric between neighbor pixels in RGB-D space. The proposed method that utilizes RGB-D hybrid constraints generates stereo segments with accuracies higher than those obtained by state-of-the-art methods. Wei Ma 0008, Luwei Yang, Shibiao Xu, Xiaopeng Zhang 0001 |
IEEE Signal Process. Lett. | 5 |
| 2016 | Speeding Up the Bilateral Filter: A Joint Acceleration WayabstractComputational complexity of the brute-force implementation of the bilateral filter (BF) depends on its filter kernel size. To achieve the constant-time BF whose complexity is irrelevant to the kernel size, many techniques have been proposed, such as 2D box filtering, dimension promotion, and shiftability property. Although each of the above techniques suffers from accuracy and efficiency problems, previous algorithm designers were used to take only one of them to assemble fast implementations due to the hardness of combining them together. Hence, no joint exploitation of these techniques has been proposed to construct a new cutting edge implementation that solves these problems. Jointly employing five techniques: kernel truncation, best N-term approximation as well as previous 2D box filtering, dimension promotion, and shiftability property, we propose a unified framework to transform BF with arbitrary spatial and range kernels into a set of 3D box filters that can be computed in linear time. To the best of our knowledge, our algorithm is the first method that can integrate all these acceleration techniques and, therefore, can draw upon one another's strong point to overcome deficiencies. The strength of our method has been corroborated by several carefully designed experiments. In particular, the filtering accuracy is significantly improved without sacrificing the efficiency at running time. Longquan Dai, Mengke Yuan, Xiaopeng Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2016 | Image Retargeting by Texture-Aware SynthesisabstractReal-world images usually contain vivid contents and rich textural details, which will complicate the manipulation on them. In this paper, we design a new framework based on exampled-based texture synthesis to enhance content-aware image retargeting. By detecting the textural regions in an image, the textural image content can be synthesized rather than simply distorted or cropped. This method enables the manipulation of textural & non-textural regions with different strategies since they have different natures. We propose to retarget the textural regions by example-based synthesis and non-textural regions by fast multi-operator. To achieve practical retargeting applications for general images, we develop an automatic and fast texture detection method that can detect multiple disjoint textural regions. We adjust the saliency of the image according to the features of the textural regions. To validate the proposed method, comparisons with state-of-the-art image retargeting techniques and a user study were conducted. Convincing visual results are shown to demonstrate the effectiveness of the proposed method. Weiming Dong, Fuzhang Wu, Yan Kong, Xing Mei, Tong-Yee Lee, Xiaopeng Zhang 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2016 | Automatic Constraint Detection for 2D Layout RegularizationabstractIn this paper, we address the problem of constraint detection for layout regularization. The layout we consider is a set of two-dimensional elements where each element is represented by its bounding box. Layout regularization is important in digitizing plans or images, such as floor plans and facade images, and in the improvement of user-created contents, such as architectural drawings and slide layouts. To regularize a layout, we aim to improve the input by detecting and subsequently enforcing alignment, size, and distance constraints between layout elements. Similar to previous work, we formulate layout regularization as a quadratic programming problem. In addition, we propose a novel optimization algorithm that automatically detects constraints. We evaluate the proposed framework using a variety of input layouts from different applications. Our results demonstrate that our method has superior performance to the state of the art. Haiyong Jiang, Liangliang Nan, Dong-Ming Yan 0001, Weiming Dong, Xiaopeng Zhang 0001, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2016 | Measuring and Predicting Visual Importance of Similar ObjectsabstractSimilar objects are ubiquitous and abundant in both natural and artificial scenes. Determining the visual importance of several similar objects in a complex photograph is a challenge for image understanding algorithms. This study aims to define the importance of similar objects in an image and to develop a method that can select the most important instances for an input image from multiple similar objects. This task is challenging because multiple objects must be compared without adequate semantic information. This challenge is addressed by building an image database and designing an interactive system to measure object importance from human observers. This ground truth is used to define a range of features related to the visual importance of similar objects. Then, these features are used in learning-to-rank and random forest to rank similar objects in an image. Importance predictions were validated on 5,922 objects. The most important objects can be identified automatically. The factors related to composition (e.g., size, location, and overlap) are particularly informative, although clarity and color contrast are also important. We demonstrate the usefulness of similar object importance on various applications, including image retargeting, image compression, image re-attentionizing, image admixture, and manipulation of blindness images. Yan Kong, Weiming Dong, Xing Mei, Chongyang Ma, Tong-Yee Lee, Siwei Lyu, Feiyue Huang, Xiaopeng Zhang 0001 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2016 | Feature-aware natural texture synthesis
Fuzhang Wu, Weiming Dong, Yan Kong, Xing Mei, Dong-Ming Yan 0001, Xiaopeng Zhang 0001, Jean-Claude Paul |
Vis. Comput. | 6 |
| 2015 | Facade Layout SymmetrizationabstractWe present an automatic algorithm for symmetrizing facade layouts. Our method symmetrizes a given facade layout while minimally modifying the original layout. Based on the principles of symmetry in urban design, we formulate the problem of facade layout symmetrization as an optimization problem. Our system further enhances the regularity of the final layout by redistributing and aligning boxes in the layout. We demonstrate that the proposed solution can generate symmetric facade layouts efficiently. Haiyong Jiang, Weiming Dong, Dong-Ming Yan 0001, Xiaopeng Zhang 0001 |
CAD/Graphics | 4 |
| 2015 | High Resolution Photography with an RGB-Infrared CameraabstractA convenient solution to RGB-Infrared photography is to extend the basic RGB mosaic with a fourth filter type with high transmittance in the near-infrared band. Unfortunately, applying conventional demosaicing algorithms to RGB-IR sensors is not possible for two reasons. First, the RGB and near-infrared image are differently focused due to different refractive indices of each band. Second, manufacturing constraints introduce crosstalk between RGB and IR channels. In this paper we propose a novel image formation model for RGB-IR cameras that can be easily calibrated, and propose an efficient algorithm that jointly addresses three restoration problems--channel deblurring, channel separation and pixel demosaicing--using quadratic image regularizers. We also extend our algorithm to handle more general regularizers and pixel saturation. Experiments show that our method produces sharp, full-resolution images of pure RGB color and IR. Huixuan Tang, Xiaopeng Zhang 0001, Shaojie Zhuo, Kiriakos N. Kutulakos, Liang Shen 0007 |
ICCP | 2 |
| 2015 | Fully Connected Guided Image FilteringabstractThis paper presents a linear time fully connected guided filter by introducing the minimum spanning tree (MST) to the guided filter (GF). Since the intensity based filtering kernel of GF is apt to overly smooth edges and the fixed-shape local box support region adopted by GF is not geometric-adaptive, our filter introduces an extra spatial term, the tree similarity, to the filtering kernel of GF and substitutes the box window with the implicit support region by establishing all-pairs-connections among pixels in the image and assigning the spatial-intensity-aware similarity to these connections. The adaptive implicit support region composed by the pixels with large kernel weights in the entire image domain has a big advantage over the predefined local box window in presenting the structure of an image for the reason that: 1, MST can efficiently present the structure of an image, 2, the kernel weight of our filter considers the tree distance defined on the MST. Due to these reasons, our filter achieves better edge-preserving results. We demonstrate the strength of the proposed filter in several applications. Experimental results show that our method produces better results than state-of-the-art methods. Longquan Dai, Mengke Yuan, Feihu Zhang, Xiaopeng Zhang 0001 |
ICCV | 4 |
| 2015 | Segment Graph Based Image Filtering: Fast Structure-Preserving SmoothingabstractIn this paper, we design a new edge-aware structure, named segment graph, to represent the image and we further develop a novel double weighted average image filter (SGF) based on the segment graph. In our SGF, we use the tree distance on the segment graph to define the internal weight function of the filtering kernel, which enables the filter to smooth out high-contrast details and textures while preserving major image structures very well. While for the external weight function, we introduce a user specified smoothing window to balance the smoothing effects from each node of the segment graph. Moreover, we also set a threshold to adjust the edge-preserving performance. These advantages make the SGF more flexible in various applications and overcome the "halo" and "leak" problems appearing in most of the state-of-the-art approaches. Finally and importantly, we develop a linear algorithm for the implementation of our SGF, which has an O(N) time complexity for both gray-scale and high dimensional images, regardless of the kernel size and the intensity range. Typically, as one of the fastest edge-preserving filters, our CPU implementation achieves 0.15s per megapixel when performing filtering for 3-channel color images. The strength of the proposed filter is demonstrated by various applications, including stereo matching, optical flow, joint depth map upsampling, edge-preserving smoothing, edges detection, image abstraction and texture editing. Feihu Zhang, Longquan Dai, Shiming Xiang, Xiaopeng Zhang 0001 |
ICCV | 4 |
| 2015 | Efficient maximal Poisson-disk sampling and remeshing on surfaces
Jianwei Guo 0003, Dong-Ming Yan 0001, Xiaohong Jia 0001, Xiaopeng Zhang 0001 |
Comput. Graph. | 4 |
| 2015 | Depth map upsampling using compressive sensing based model
Longquan Dai, Haoxing Wang, Xiaopeng Zhang 0001 |
Neurocomputing | 3 |
| 2015 | A Survey of Blue-Noise Sampling and Its Applications
Dong-Ming Yan 0001, Jianwei Guo 0003, Bin Wang 0021, Xiaopeng Zhang 0001, Peter Wonka |
J. Comput. Sci. Technol. | 4 |
| 2015 | 3D shape retrieval using viewpoint information-theoretic measuresabstractAbstract In this paper, we present an information‐theoretic framework to compute the shape similarity between 3D polygonal models. Given a 3D model, an information channel between a sphere of viewpoints around the model and its polygonal mesh is defined to compute the specific information associated with each viewpoint. The obtained information sphere can be seen as a shape descriptor of the model. Then, given two models, their similarity is obtained by performing a registration process between the corresponding information spheres. The distance between the information histograms is also defined as a coarse measure of similarity, as well as the scalar value given by the mutual information of the channel. The performance of all these measures is tested using the Princeton Shape Benchmark database. Copyright © 2013 John Wiley & Sons, Ltd. Xavier Bonaventura, Jianwei Guo 0003, Weiliang Meng, Miquel Feixas, Xiaopeng Zhang 0001, Mateu Sbert |
Comput. Animat. Virtual Worlds | 5 |
| 2015 | Fast Minimax Path-Based Joint Depth InterpolationabstractWe propose a fast minimax path-based depth interpolation method. The algorithm computes for each target pixel varying contributions from reliable depth seeds, and weighted averaging is used to interpolate missing depths. Compared with state-of-the-art joint geodesic upsampling method which selects the K nearest seeds to interpolate missing depths with O(Kn) complexity, our method does not need to limit the number of seeds to K and reduces the computational complexity to O(n). In addition, the minimax path chooses a path with the smallest maximum immediate pairwise pixel difference on it, so it tends to preserve sharp depth discontinuities better. In contrast to the results of previous depth upsampling algorithms, our approach can provide accurate depths with fewer artifacts. Longquan Dai, Feihu Zhang, Xing Mei, Xiaopeng Zhang 0001 |
IEEE Signal Process. Lett. | 4 |
| 2015 | Robust Rooftop Extraction From Visible Band Images Using Higher Order CRFabstractIn this paper, we propose a robust framework for building extraction in visible band images. We first get an initial classification of the pixels based on an unsupervised presegmentation. Then, we develop a novel conditional random field (CRF) formulation to achieve accurate rooftops extraction, which incorporates pixel-level information and segment-level information for the identification of rooftops. Comparing with the commonly used CRF model, a higher order potential defined on segment is added in our model, by exploiting region consistency and shape feature at segment level. Our experiments show that the proposed higher order CRF model outperforms the state-of-the-art methods both at pixel and object levels on rooftops with complex structures and sizes in challenging environments. Er Li, John Femiani 0001, Shibiao Xu, Xiaopeng Zhang 0001, Peter Wonka |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2015 | PM-PM: PatchMatch With Potts Model for Object Segmentation and Stereo MatchingabstractThis paper presents a unified variational formulation for joint object segmentation and stereo matching, which takes both accuracy and efficiency into account. In our approach, depth-map consists of compact objects, each object is represented through three different aspects: 1) the perimeter in image space; 2) the slanted object depth plane; and 3) the planar bias, which is to add an additional level of detail on top of each object plane in order to model depth variations within an object. Compared with traditional high quality solving methods in low level, we use a convex formulation of the multilabel Potts Model with PatchMatch stereo techniques to generate depth-map at each image in object level and show that accurate multiple view reconstruction can be achieved with our formulation by means of induced homography without discretization or staircasing artifacts. Our model is formulated as an energy minimization that is optimized via a fast primal-dual algorithm, which can handle several hundred object depth segments efficiently. Performance evaluations in the Middlebury benchmark data sets show that our method outperforms the traditional integer-valued disparity strategy as well as the original PatchMatch algorithm and its variants in subpixel accurate disparity estimation. The proposed algorithm is also evaluated and shown to produce consistently good results for various real-world data sets (KITTI benchmark data sets and multiview benchmark data sets). Shibiao Xu, Feihu Zhang, Xiaofei He 0001, Xukun Shen, Xiaopeng Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2014 | Blue-Noise Remeshing with Farthest Point OptimizationabstractAbstract In this paper, we present a novel method for surface sampling and remeshing with good blue‐noise properties. Our approach is based on the farthest point optimization (FPO), a relaxation technique that generates high quality blue‐noise point sets in 2D. We propose two important generalizations of the original FPO framework: adaptive sampling and sampling on surfaces. A simple and efficient algorithm for accelerating the FPO framework is also proposed. Experimental results show that the generalized FPO generates point sets with excellent blue‐noise properties for adaptive and surface sampling. Furthermore, we demonstrate that our remeshing quality is superior to the current state‐of‐theߚart approaches. Dong-Ming Yan 0001, Jianwei Guo 0003, Xiaohong Jia 0001, Xiaopeng Zhang 0001, Peter Wonka |
Comput. Graph. Forum | 4 |
| 2014 | Inverse procedural modeling of facade layoutsabstractIn this paper, we address the following research problem: How can we generate a meaningful split grammar that explains a given facade layout? To evaluate if a grammar is meaningful, we propose a cost function based on the description length and minimize this cost using an approximate dynamic programming framework. Our evaluation indicates that our framework extracts meaningful split grammars that are competitive with those of expert users, while some users and all competing automatic solutions are less successful. Fuzhang Wu, Dong-Ming Yan 0001, Weiming Dong, Xiaopeng Zhang 0001, Peter Wonka |
ACM Trans. Graph. | 4 |
| 2014 | Summarization-Based Image Resizing by Intelligent Object CarvingabstractImage resizing can be more effectively achieved with a better understanding of image semantics. In this paper, similar patterns that exist in many real-world images are analyzed. By interactively detecting similar objects in an image, the image content can be summarized rather than simply distorted or cropped. This method enables the manipulation of image pixels or patches as well as semantic objects in the scene during image resizing process. Given the special nature of similar objects in a general image, the integration of a novel object carving (OC) operator with the multi-operator framework is proposed for summarizing similar objects. The object removal sequence in the summarization strategy directly affects resizing quality. The method by which to evaluate the visual importance of the object as well as to optimally select the candidates for object carving is demonstrated. To achieve practical resizing applications for general images, a template matching-based method is developed. This method can detect similar objects even when they are of various colors, transformed in terms of perspective, or partially occluded. To validate the proposed method, comparisons with state-of-the-art resizing techniques and a user study were conducted. Convincing visual results are shown to demonstrate the effectiveness of the proposed method. Weiming Dong, Tong-Yee Lee, Fuzhang Wu, Yan Kong, Xiaopeng Zhang 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2014 | Low-Resolution Remeshing Using the Localized Restricted Voronoi DiagramabstractA big problem in triangular remeshing is to generate meshes when the triangle size approaches the feature size in the mesh. The main obstacle for Centroidal Voronoi Tessellation (CVT)-based remeshing is to compute a suitable Voronoi diagram. In this paper, we introduce the localized restricted Voronoi diagram (LRVD) on mesh surfaces. The LRVD is an extension of the restricted Voronoi diagram (RVD), but it addresses the problem that the RVD can contain Voronoi regions that consist of multiple disjoint surface patches. Our definition ensures that each Voronoi cell in the LRVD is a single connected region. We show that the LRVD is a useful extension to improve several existing mesh-processing techniques, most importantly surface remeshing with a low number of vertices. While the LRVD and RVD are identical in most simple configurations, the LRVD is essential when sampling a mesh with a small number of points and for sampling surface areas that are in close proximity to other surface areas, e.g., nearby sheets. To compute the LRVD, we combine local discrete clustering with a global exact computation. Dong-Ming Yan 0001, Guanbo Bao, Xiaopeng Zhang 0001, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2014 | Data-Driven Synthetic Modeling of TreesabstractIn this paper, we develop a data-driven technique to model trees from a single laser scan. A multi-layer representation of the tree structure is proposed to guide the modeling process. In this process, a marching cylinder algorithm is first developed to construct visible branches from the laser scan data. Three levels of crown feature points are then extracted from the scan data to synthesize three layers of non-visible branches. Based on the hierarchical particle flow technique, the branch synthesis method has the advantage of producing visually convincing tree models that are consistent with scan data. User intervention is extremely limited. The robustness of this technique has been validated on both conifer and broadleaf trees. Xiaopeng Zhang 0001, Hongjun Li 0002, Mingrui Dai, Wei Ma 0008, Long Quan |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2014 | Efficient triangulation of Poisson-disk sampled point sets
Jianwei Guo 0003, Dong-Ming Yan 0001, Guanbo Bao, Weiming Dong, Xiaopeng Zhang 0001, Peter Wonka |
Vis. Comput. | 5 |
| 2013 | Segment-Tree Based Cost Aggregation for Stereo MatchingabstractThis paper presents a novel tree-based cost aggregation method for dense stereo matching. Instead of employing the minimum spanning tree (MST) and its variants, a new tree structure, "Segment-Tree", is proposed for non-local matching cost aggregation. Conceptually, the segment-tree is constructed in a three-step process: first, the pixels are grouped into a set of segments with the reference color or intensity image, second, a tree graph is created for each segment, and in the final step, these independent segment graphs are linked to form the segment-tree structure. In practice, this tree can be efficiently built in time nearly linear to the number of the image pixels. Compared to MST where the graph connectivity is determined with local edge weights, our method introduces some 'non-local' decision rules: the pixels in one perceptually consistent segment are more likely to share similar disparities, and therefore their connectivity within the segment should be first enforced in the tree construction process. The matching costs are then aggregated over the tree within two passes. Performance evaluation on 19 Middlebury data sets shows that the proposed method is comparable to previous state-of-the-art aggregation methods in disparity accuracy and processing speed. Furthermore, the tree structure can be refined with the estimated disparities, which leads to consistent scene segmentation and significantly better aggregation results. Xing Mei, Weiming Dong, Haitao Wang 0006, Xiaopeng Zhang 0001 |
CVPR | 5 |
| 2013 | Cross-Field Joint Image Restoration via Scale MapabstractColor, infrared, and flash images captured in different fields can be employed to effectively eliminate noise and other visual artifacts. We propose a two-image restoration framework considering input images in different fields, for example, one noisy color image and one dark-flashed near infrared image. The major issue in such a framework is to handle structure divergence and find commonly usable edges and smooth transition for visually compelling image reconstruction. We introduce a scale map as a competent representation to explicitly model derivative-level confidence and propose new functions and a numerical solver to effectively infer it following new structural observations. Our method is general and shows a principled way for cross-field restoration. Qiong Yan, Xiaoyong Shen, Li Xu 0001, Shaojie Zhuo, Xiaopeng Zhang 0001, Liang Shen 0007, Jiaya Jia |
ICCV | 5 |
| 2013 | Near-infrared guided color image dehazingabstractNear-infrared (NIR) light has stronger penetration capability than visible light due to its long wavelengths and is thus less scattered by particles in the air. This makes it desirable for image dehazing to unveil details of distant objects in landscape photographs. In this paper, we propose an improved image dehazing scheme using a pair of color and NIR images, which effectively estimates the airlight color and transfers details from the NIR. A two-stage dehazing method is proposed by exploiting the dissimilarity between RGB and NIR for airlight color estimation, followed by a dehazing procedure through an optimization framework. Experiments on captured haze images show that our method can achieve substantial improvements on the detail recovery and the color distribution over the existing image dehazing algorithms. Shaojie Zhuo, Xiaopeng Zhang 0001, Liang Shen 0007, Sabine Süsstrunk |
ICIP | 3 |
| 2013 | Illustrating the disassembly of 3D models
Jianwei Guo 0003, Dong-Ming Yan 0001, Er Li, Weiming Dong, Peter Wonka, Xiaopeng Zhang 0001 |
Comput. Graph. | 6 |
| 2013 | Content-Based Colour TransferabstractAbstract This paper presents a novel content‐based method for transferring the colour patterns between images. Unlike previous methods that rely on image colour statistics, our method puts an emphasis on high‐level scene content analysis. We first automatically extract the foreground subject areas and background scene layout from the scene. The semantic correspondences of the regions between source and target images are established. In the second step, the source image is re‐coloured in a novel optimization framework, which incorporates the extracted content information and the spatial distributions of the target colour styles. A new progressive transfer scheme is proposed to integrate the advantages of both global and local transfer algorithms, as well as avoid the over‐segmentation artefact in the result. Experiments show that with a better understanding of the scene contents, our method well preserves the spatial layout, the colour distribution and the visual coherence in the transfer process. As an interesting extension, our method can also be used to re‐colour video clips with spatially‐varied colour effects. Fuzhang Wu, Weiming Dong, Yan Kong, Xing Mei, Jean-Claude Paul, Xiaopeng Zhang 0001 |
Comput. Graph. Forum | 6 |
| 2013 | Ricci flow-based spherical parameterization and surface registration
Huiguang He, Guangyu Zou, Xiaopeng Zhang 0001, Xianfeng Gu, Jing Hua 0001 |
Comput. Vis. Image Underst. | 4 |
| 2013 | Statistical learning based facial animationabstractTo synthesize real-time and realistic facial animation, we present an effective algorithm which combines image- and geometry-based methods for facial animation simulation. Considering the numerous motion units in the expression coding system, we present a novel simplified motion unit based on the basic facial expression, and construct the corresponding basic action for a head model. As image features are difficult to obtain using the performance driven method, we develop an automatic image feature recognition method based on statistical learning, and an expression image semi-automatic labeling method with rotation invariant face detection, which can improve the accuracy and efficiency of expression feature identification and training. After facial animation redirection, each basic action weight needs to be computed and mapped automatically. We apply the blend shape method to construct and train the corresponding expression database according to each basic action, and adopt the least squares method to compute the corresponding control parameters for facial animation. Moreover, there is a pre-integration of diffuse light distribution and specular light distribution based on the physical method, to improve the plausibility and efficiency of facial rendering. Our work provides a simplification of the facial motion unit, an optimization of the statistical training process and recognition process for facial animation, solves the expression parameters, and simulates the subsurface scattering effect in real time. Experimental results indicate that our method is effective and efficient, and suitable for computer animation and interactive applications. Shibiao Xu, Guanghui Ma, Weiliang Meng, Xiaopeng Zhang 0001 |
J. Zhejiang Univ. Sci. C | 4 |
| 2013 | SimLocator: robust locator of similar objects in images
Yan Kong, Weiming Dong, Xing Mei, Xiaopeng Zhang 0001, Jean-Claude Paul |
Vis. Comput. | 4 |
| 2012 | Large-scale forest rendering: Real-time, realistic, and progressive
Guanbo Bao, Hongjun Li 0002, Xiaopeng Zhang 0001, Weiming Dong |
Comput. Graph. | 3 |
| 2012 | Real-time ink simulation using a grid-particle method
Shibiao Xu, Xing Mei, Weiming Dong, Xiaopeng Zhang 0001 |
Comput. Graph. | 5 |
| 2012 | Easy modeling of realistic trees from freehand sketches
Zhiguo Jiang 0001, Hongjun Li 0002, Xiaopeng Zhang 0001 |
Frontiers Comput. Sci. | 4 |
| 2012 | Fast Multi-Operator Image Resizing and Evaluation
Weiming Dong, Guanbo Bao, Xiaopeng Zhang 0001, Jean-Claude Paul |
J. Comput. Sci. Technol. | 3 |
| 2011 | Hardware instancing for real-time realistic forest renderingabstractReal-time rendering of vegetation is important in many applications, such as video games, internet graphics applications, landscape design and visualization. However, the visualization of large-scale forests has always been a challenge not only due to the high geometric complexity but also due to the small batch problem. Moreover, generating real-time shadows for forests will heavily increase the burden. The batch problem is caused by a large number of graphics API draw calls launched in every frame. Normally, rendering a tree model requires at least one graphics API draw call. As a forest usually consists of thousands of trees and the graphics API invocation is a relatively high cost for CPU, the large-scale forest rendering is often CPU bound. Guanbo Bao, Weiliang Meng, Hongjun Li 0002, Xiaopeng Zhang 0001 |
SIGGRAPH Asia Sketches | 5 |
| 2011 | Translucent material transfer based on single imagesabstractExtraction and re-rendering of real materials give large contributions to various image-based applications. As one of the key properties of modeling the appearance of an object, materials mainly focus on the effects caused by light transportation. Therefore, understanding the characteristics of a complex material from a single photograph and transferring it to an object in another image becomes a very challenging problem. Weiming Dong, Xiaopeng Zhang 0001, Jean-Claude Paul |
SIGGRAPH Asia Sketches | 4 |
| 2011 | Distribution-aware image color transferabstractColor transfer is a practical image editing technology which is useful in various applications. An ideal color transfer algorithm should keep the scene in the source image and apply the color styles of the reference image. All the dominant color styles of the reference image should be presented in the result especially when there are similar contents in the source and reference images. Fuzhang Wu, Weiming Dong, Xing Mei, Xiaopeng Zhang 0001, Xiaohong Jia 0001, Jean-Claude Paul |
SIGGRAPH Asia Sketches | 4 |
| 2011 | Ridge extraction of a smooth 2-manifold surface based on vector field
Wujun Che, Xiaopeng Zhang 0001, Yi-Kuan Zhang, Jean-Claude Paul, Bo Xu 0002 |
Comput. Aided Geom. Des. | 2 |
| 2011 | Direct quad-dominant meshing of point cloud via global parameterization
Er Li, Wujun Che, Xiaopeng Zhang 0001, Yi-Kuan Zhang, Bo Xu 0002 |
Comput. Graph. | 3 |
| 2011 | Meshless quadrangulation by global parameterization
Er Li, Bruno Lévy 0001, Xiaopeng Zhang 0001, Wujun Che, Weiming Dong, Jean-Claude Paul |
Comput. Graph. | 3 |
| 2010 | Enhancing low light images using near infrared flash imagesabstractIn low light environment, photographs taken with a high ISO setting suffer from significant noise. In this paper, we propose to use a near infrared (NIR) flash image, instead of a normal visible flash image, to enhance its corresponding noisy visible image. We build a hybrid camera system to take an visible image and its NIR counterpart simultaneously. We introduce a new method to denoise an visible image and enhance its details using its corresponding NIR flash image. Experimental results show the superiority of our method compared with previous image denoising and detail enhancement methods. Shaojie Zhuo, Xiaopeng Zhang 0001, Xiaoping Miao, Terence Sim |
ICIP | 2 |
| 2010 | Fast local color transfer via dominant colors mappingabstractColor transfer is an image editing technique which arises various applications, from daily photo appearance enhancement to movie post-processing. An ideal color transfer algorithm should keep the scene from the source image and apply the color style of the target image. All the dominant colors in the target image should be transferred to the source, while the colors in the source image which are apparently distinct from the target style should not appear in the result. The preservation of the scene details is also important for a good color transfer algorithm. Weiming Dong, Guanbo Bao, Xiaopeng Zhang 0001, Jean-Claude Paul |
SIGGRAPH ASIA (Sketches) | 3 |
| 2010 | Automatic architecture model generation based on object hierarchyabstractTerrestrial laser scanner (TLS) can be used to acquire 3D facade information of modern architectures, represented as point cloud data (PCD). Basic shape elements of an architecture, like windows and doors, should be recovered in reconstruction; and the model should be represented corresponding to the information of architectural design, such as lines and polygons. Most recent approaches could not reconstruct models automatically with designed shape details. Either user's interactions are needed [Zheng et al. 2010; Nan et al. 2010]; or the reconstructed model is coarse without information of shape details. Therefore, it is necessary to develop new algorithms to generate geometric models automatically, fitting well the design information of architectural PCD. A novel framework is proposed to generate explicitly an architectural model from scanned points of an existing architecture. An automatic, hierarchical and fast facade reconstruction framework is presented based on a novel combination of facade structures, detailed windows propagation, hierarchical model consolidation and contextual semantic representations. As a result, a high-quality geometric model of an architecture ia generated. Figure 1 shows the procedure of this work, from building detection, to planar region decomposition, to boundary point extraction, and to the consolidated hierarchal model. Xiaojuan Ning, Xiaopeng Zhang 0001, Yinghui Wang 0001 |
SIGGRAPH ASIA (Sketches) | 2 |
| 2010 | Multiresolution foliage for forest renderingabstractAbstract Plants are important objects in virtual environments. High complexity of shape structure is found in plant communities. Level of detail (LOD) of plant geometric models becomes important for interactive forest rendering. We emphasize three major problems in current research: the time consumption in LOD model construction and extraction, the balance between visual effect and data compression, and the time consumption in the communication between Central Processing Unit (CPU) and Graphics Processing Unit (GPU). We present a new foliage simplification framework for LOD model and forest rendering. By an uneven subdivision of the tree crown volume, the cost for LOD model construction is drastically reduced. With a GPU‐oriented design of LOD storage structure for foliage, the costly hierarchical traversal of a binary tree is replaced by a sequential lookup of an array. The structure also decreases the communication between the CPU and the GPU in rendering. In addition, Leaf density is introduced to adapt compression to the local distribution of leaves, so that more visually relevant details are kept. According to foliage nature (broad leaves or needles), higher compression are finally reached using mixed polygon/line models. This framework is implemented on virtual scenes of simulated trees with high detail. Copyright © 2009 John Wiley & Sons, Ltd. Qingqiong Deng, Xiaopeng Zhang 0001, Marc Jaeger 0002 |
Comput. Animat. Virtual Worlds | 2 |
| 2009 | Computing lines of curvature for implicit surfaces
Xiaopeng Zhang 0001, Wujun Che, Jean-Claude Paul |
Comput. Aided Geom. Des. | 1 |
| 2009 | Estimating differential quantities from point cloud based on a linear fitting of normal vectors
Zhanglin Cheng, Xiaopeng Zhang 0001 |
Sci. China Ser. F Inf. Sci. | 2 |
| 2009 | Modeling plants with sensor data
Wei Ma 0008, Bo Xiang, Hongbin Zha, Xiaopeng Zhang 0001 |
Sci. China Ser. F Inf. Sci. | 5 |
| 2009 | Optimized image resizing using seam carving and scalingabstractWe present a novel method for content-aware image resizing based on optimization of a well-defined image distance function, which preserves both the important regions and the global visual effect (the background or other decorative objects) of an image. The method operates by joint use of seam carving and image scaling. The principle behind our method is the use of a bidirectional similarity function of image Euclidean distance (IMED), while cooperating with a dominant color descriptor (DCD) similarity and seam energy variation. The function is suitable for the quantitative evaluation of the resizing result and the determination of the best seam carving number. Different from the previous simplex-mode approaches, our method takes the advantages of both discrete and continuous methods. The technique is useful in image resizing for both reduction/retargeting and enlarging. We also show that this approach can be extended to indirect image resizing. Weiming Dong, Jean-Claude Paul, Xiaopeng Zhang 0001 |
ACM Trans. Graph. | 4 |
| 2008 | Enhancing photographs with Near Infra-Red imagesabstractNear Infra-Red (NIR) images of natural scenes usually have better contrast and contain rich texture details that may not be perceived in visible light photographs (VIS). In this paper, we propose a novel method to enhance a photograph by using the contrast and texture information of its corresponding NIR image. More precisely, we first decompose the NIR/VIS pair into average and detail wavelet subbands. We then transfer the contrast in the average subband and transfer texture in the detail subbands. We built a special camera mount that optically aligns two consumer-grade digital cameras, one of which was modified to capture NIR. Our results exhibit higher visual quality than tone-mapped HDR images, showing that NIR imaging is useful for computational photography. Xiaopeng Zhang 0001, Terence Sim, Xiaoping Miao |
CVPR | 1 |
| 2008 | Objective Evaluation of 3D Reconstructed Plants and Trees from 2D ImagesabstractIn this work, we propose a preliminary study on objective approaches in evaluating the synthesized 3D plants from the given image data. Two measures, in a general name of "resemblance index (RI)", are defined and investigated based on the normalized mutual information (RINI), and coincident-bit-counting criterion (RICBC), respectively. We propose a strategy of hierarchical evaluations on different attributes of objects. In this study, we only consider the geometrical and crown distribution attributes. The numerical investigations confirm the approach to be useful for a fast evaluation of 3D reconstructed plants/trees from image data. Discussions are given about the two measures based on the numerical examples. Bao-Gang Hu, Xiaopeng Zhang 0001, Marc Jaeger 0002 |
CW | 2 |
| 2008 | Real-Time Marker Level Set on GPUabstractLevel set methods have been extensively used to track the dynamical interfaces between different materials for physically based simulation, geometry modeling, oceanic modeling and other scientific and engineering applications. Due to the inherent Eulerian characteristics, interface evolution based on level set usually suffers from numerical diffusion, sharp feature missing and mass loss. Although some effective methods such as Particle Level Set (PLS) and Marker Level Set (MLS) have been proposed to tackle these difficulties, the complicated correction process and the high computational cost pose severe limitations for real-time applications. In this paper we provide an efficient parallel implementation of the Marker Level Set method on latest graphics hardware. Each step of the MLS method is fully mapped on GPU with an innovative combination of different computation techniques. Relying on GPU's parallelism and flexible programmability, the method provides real-time performance for large size 2D examples and moderate 3D examples, which is significantly faster than previous CPU-based methods. Xing Mei, Philippe Decaudin, Bao-Gang Hu, Xiaopeng Zhang 0001 |
CW | 4 |
| 2008 | Decomposition of branching volume data by tip detectionabstractWe present an approach to decomposing branching volume data into sub-branches. First, a metric is proposed for evaluating local convexities in volumetric data, and it is a criterion for global selection of tip points. Second, a multi-path growing strategy is adopted to segment the volumes based on a DFS transformation starting from the tips. Experiments show that this approach is capable of generating desirable components and reasonable segmentation boundaries of a volume. Wei Ma 0008, Bo Xiang, Xiaopeng Zhang 0001, Hongbin Zha |
ICIP | 3 |
| 2008 | Image-based plant modeling by knowing leaves from their apexesabstractIn the paper, we present a novel approach to modeling plants from images by detecting apex features. First, an effective algorithm is proposed to extract apex features in volumetric data recovered from the images. It provides position and pose information for assigning 3D generic leaves. Then, the 3D leaf shapes are determined by an optimization based on the volume. Finally, Branches are modeled by using a particle flow approach. The proposed method is simply with limited manual intervention and has the obvious benefit of knowing a leaf by its visible apex part. Wei Ma 0008, Hongbin Zha, Xiaopeng Zhang 0001, Bo Xiang |
ICPR | 4 |
| 2007 | Lines of curvature and umbilical points for implicit surfaces
Wujun Che, Jean-Claude Paul, Xiaopeng Zhang 0001 |
Comput. Aided Geom. Des. | 3 |
| 2007 | Simple Reconstruction of Tree Branches from a Single Range Image
Zhanglin Cheng, Xiaopeng Zhang 0001, Baoquan Chen |
J. Comput. Sci. Technol. | 2 |
| 2006 | Rain Removal in Video by Combining Temporal and Chromatic PropertiesabstractRemoval of rain streaks in video is a challenging problem due to the random spatial distribution and fast motion of rain. This paper presents a new rain removal algorithm that incorporates both temporal and chromatic properties of rain in video. The temporal property states that an image pixel is never always covered by rain throughout the entire video. The chromatic property states that the changes of R, G, and B values of rainaffected pixels are approximately the same. By using both properties, the algorithm can detect and remove rain streaks in both stationary and dynamic scenes taken by stationary cameras. To handle videos taken by moving cameras, the video can be stabilized for rain removal, and destabilized to restore camera motion after rain removal. It can handle both light rain and heavy rain conditions. Experimental results show that the algorithm performs better than existing algorithms. Xiaopeng Zhang 0001, Hao Li 0032, Yingyi Qi, Wee Kheng Leow, Teck Khim Ng |
ICME | 1 |
| 2005 | Enclosure Sphere Based Cell Visibility for Virtual Endoscopy
Xiaopeng Zhang 0001 |
IEEE Visualization | 2 |
| 2001 | Hair Image Generation Using Connected Texels
Xiaopeng Zhang 0001, Yanyun Chen, Enhua Wu |
J. Comput. Sci. Technol. | 1 |