Yuanfeng Zhou

dblp:97/2964 · DBLP profile ↗
← Back
105ranked-venue papers
4as first author
75since 2021 · last 2026
0000-0001-6950-3261ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 80 · 4 first-author · 52 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 16 since 2021Artificial intelligence and machine learning · 12 · 10 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DentalGS: Pose-Free 3D Gaussian Splatting from Five Intraoral Images for Novel View Synthesis
abstract
Orthodontic treatment needs regular tooth alignment checks, but current methods depend on clinic visits, limiting remote care. With the emergence of 3D Gaussian Splatting (3DGS), realistic novel views can be synthesized, making it possible for clinicians to remotely monitor orthodontic conditions. However, using only five intraoral images with unknown camera poses and dynamic lighting presents major challenges in dental applications. To address these challenges, we propose DentalGS, an enhanced 3DGS framework capable of synthesizing novel intraoral views from five post-orthodontic intraoral images and pre-orthodontic intraoral scan (IOS) data as prior, without camera poses. Our method initializes a Gaussian point cloud labeled with ISO-FDI tooth classes based on the patient’s pre-orthodontic IOS data, then estimates camera poses through iterative optimization. We introduce a Progressive Pair Generation Strategy as a data augmentation method that generates damage–repair image pairs to train a RepairNet, aiming to restore degraded geometry and appearance caused by the limited number of intraoral images. Additionally, we introduce a Lighting-Aware 3DGS inspired by physical reflectance properties to mitigate the effects of dynamic lighting conditions. Experimental results show that our method produces high-quality novel views while preserving geometric structure even under extreme viewpoints, offering an efficient and reliable solution for 3D tooth visualization in remote orthodontic monitoring.
Honghao Dai, Yuanfeng Zhou, Guangshun Wei, Wenping Wang 0001
AAAI2
2026 Towards Event-guided Panoramic HDR Video Reconstruction for Indoor Immersive VR: A Novel Dataset and Approach
abstract
High Dynamic Range (HDR) panoramic video is crucial to enhance immersive experience in Virtual Reality (VR). However, a hurdle is that panoramic cameras often struggle with limited dynamic range and motion blur. Inspired by the event-driven sensing of the human eye, this paper explores the potential of the event cameras to enhance panoramic HDR video reconstruction. As a pioneering research endeavor, we first starts by designing a novel hybrid imaging platform equipped with preprocessing pipelines for event-panorama synchronization, alignment, and HDR ground truth generation. Based on the platform, we then introduce Ev-Pano, the first event-panorama HDR video covering diverse indoor scenes for panoramic HDR video reconstruction. We hope Ev-Pano will establish a foundation to support event-guided panoramic HDR imaging and VR research community. With Ev-Pano, we further propose a novel approach that employs a weighting function-based luminance fusion to enable events to recover missing textures in LDR panoramic videos for panoramic HDR video reconstruction. We conduct extensive experiments to demonstrate the effectiveness of our approach. The results show the best performance of ours than prior arts. Meanwhile, a user study on an HDR-capable head-mounted display (Apple Vision Pro) shows feasible perceptual quality (which is closer to the HDR ground truth) of the reconstructed panoramic HDR videos. The codes and part of the dataset can be accessed via the anonymized link https://anonymous.4open.science/r/Ev-Pano-D2D2/.
Xucheng Guo, Majed Elwardy, Yan Hu 0003, Yuanfeng Zhou, Xiaoming Chen 0006, Yiran Shen 0001
VR7
2026 Dynolayout: robust layout estimation from event-stream for extended reality under dynamic scenarios
Xucheng Guo, Qiang Qu 0004, Guangrong Zhao, Yuanfeng Zhou, Yiran Shen 0001
CCF Trans. Pervasive Comput. Interact.6
2026 HDFLStyler: Hierarchical domain-invariant feature learning for source-free domain generalization
Deqian Mao, Shanshan Gao 0003, Faqiang Huang, Caiming Zhang 0001, Yuanfeng Zhou
Neural Networks5
2026 Mean teacher based on class prototype contrast for domain adaptive object detection
Fukang Zhang, Shanshan Gao 0003, Honghao Dai, Yuanfeng Zhou
Neural Networks6
2026 Face Presentation Attack Detection by Exploiting Prior Knowledge of Region Relationships
abstract
Face recognition systems have been widely deployed in mobile devices for user authentication and payment applications. However, these biometric systems remain vulnerable to face presentation attacks, posing significant security risks. In recent years, numerous countermeasures have been proposed, with analysis of differences between bona fide and attack presentations being a commonly adopted strategy. Nevertheless, the variations in image attributes and region movements have not been thoroughly explored. In attack images, the textures of the facial region and the background tend to be more similar, while local regions often exhibit more consistent directions of movement compared to bona fide presentations. Motivated by this observation, we propose a novel face presentation attack detection method that leverages prior knowledge of region relationships. Specifically, each input face image sequence is first divided into small patches, which are then processed by a pre-trained$TimeSformer$network utilizing divided time and space attention mechanisms to extract deep features. Two metrics—$Cosine$similarity and mean squared error ($MSE$)—are subsequently employed to measure the texture similarity and movement relationships of the regions of interest. During the inference phase, these measurements are fused to distinguish bona fide from attack presentations. Extensive ablation and comparison experiments, conducted on six face presentation attack detection (PAD) databases (i.e., Idiap Replay-Attack, CASIA-MFSD, OULU-NPU, MSU-MFSD, 3DMAD, and HKBU-MARs V1+), demonstrate that our method achieves superior detection performance, significantly improving precision over state-of-the-art approaches in most experimental settings.
Lei Li 0008, Shanshan Gao 0003, Zhaoqiang Xia, Fabio Roli, Yuanfeng Zhou
IEEE Trans. Dependable Secur. Comput.5
2026 Progressive Orthodontic Motion Planning Based on Hierarchical Diffusion Transformer
abstract
Orthodontic motion planning plays a crucial role in digital orthodontics by predicting tooth motion sequences to assist dentists in formulating treatment plans efficiently. Most prior work generates the entire intermediate tooth motion sequence given the initial and target tooth alignments. In practice, only the initial alignment of the patient is obtained. However, no existing method can predict the complete motion sequence using only the initial tooth alignment. To address this gap, we propose OrthoDiff, a novel target-free framework that uses only initial tooth alignment through a progressive generation strategy. This strategy generates tooth motion sequences by decomposing the entire motion sequence into multi-level motions, progressively constraining the inference space and reducing the complexity of target-free planning from coarse to fine. Moreover, we design a hierarchical diffusion transformer as the backbone of OrthoDiff, which treats tooth alignment as a sequence of tooth tokens and fully leverages the topological prior knowledge of the dental model. Through extensive evaluations, we demonstrate that our method significantly outperforms state-of-the-art techniques in target-free tooth motion generation. Ablation studies further confirm the efficacy of key components in our network design. Meanwhile, we also achieve state-of-the-art results in tooth target alignment prediction, benefiting from our framework. The code and data will be publicly available at https://github.com/Intelligent-Orthodontics/OrthoDiff.github.io.
Yeying Fan, Yuanfeng Zhou, Guangshun Wei, Zhiming Cui 0001, Yiran Shen 0001, Yong-Jin Liu 0001, Wenping Wang 0001
IEEE Trans. Medical Imaging2
2026 SAND: Spatially Adaptive Network Depth for Fast Sampling of Neural Implicit Surfaces
abstract
Implicit neural representations are powerful for geometric modeling, but their practical use is often limited by the high computational cost of network evaluations. We observe that implicit representations require progressively lower accuracy as query points move farther from the target surface, and that even within the same iso-surface, representation difficulty varies spatially with local geometric complexity. However, conventional neural implicit models evaluate all query points with the same network depth and computational cost, ignoring this spatial variation and thereby incurring substantial computational waste. Motivated by this observation, we propose an efficient neural implicit geometry representation framework with spatially adaptive network depth (SAND). SAND leverages a volumetric network-depth map together with a tailed multi-layer perceptron (T-MLP) to model implicit representation. The volumetric depth map records, for each spatial region, the network depth required to achieve sufficient accuracy, while the T-MLP is a modified MLP designed to learn implicit functions such as signed distance functions, where an output branch, referred to as a tail, is attached to each hidden layer. This design allows network evaluation to terminate adaptively without traversing the full network and directs computational resources to geometrically important and complex regions, improving efficiency while preserving high-fidelity representations. Extensive experimental results demonstrate that our approach can significantly improve the inference-time query speed of implicit neural representations.
Chuanxiang Yang, Junhui Hou, Yuan Liu 0025, Guangshun Wei, Taku Komura, Yuanfeng Zhou, Wenping Wang 0001
ACM Trans. Graph.7
2025 ConsMatch: A Semi-Supervised Segmentation Approach for Dental CBCT by Leveraging Geometric Information to Refine Pseudo-Labels
abstract
Precise tooth instance segmentation from dental CBCT images is essential for accurate diagnosis, yet the scarcity of labeled data and the complex geometric variations of teeth make this task challenging. To address these issues, we propose ConsMatch, a semi-supervised framework that explicitly integrates geometric information into the learning process. It establishes task-level consistency between instance segmentation and boundary extraction, guiding the model to capture finegrained geometric structures. Furthermore, two geometry-aware strategies-Threshold Adjustment Strategy (TAS) and Weight Adjustment Strategy (WAS)-dynamically refine pseudo-label generation by adapting class-specific thresholds and supervision weights based on geometric consistency. This enables the model to focus on high-confidence, structure-consistent pseudolabels, enhancing training stability and segmentation accuracy. Experimental results on dental CBCT data show that ConsMatch achieves superior performance across Dice, Jaccard, and HD95 metrics, consistently outperforming existing semisupervised methods.
Shuyi Lu, Zhiming Cui 0001, Chuanxiang Yang, Guangshun Wei, Yuanfeng Zhou
BIBM7
2025 Te3DFR: Texture-Enabled 3D Face Reconstruction from Monocular Image via Self-supervised Learning
Yishen Bi, Chen Wang 0054, Lei Li 0008, Yiran Shen 0001, Yuanfeng Zhou
CGI (2)5
2025 PCDreamer: Point Cloud Completion Through Multi-view Diffusion Priors
abstract
This paper presents PCDreamer, a novel method for point cloud completion. Traditional methods typically extract features from partial point clouds to predict missing regions, but the large solution space often leads to unsatisfactory results. More recent approaches have started to use images as extra guidance, effectively improving performance, but obtaining paired data of images and partial point clouds is challenging in practice. To overcome these limitations, we harness the relatively view-consistent multi-view diffusion priors within large models, to generate novel views of the desired shape. The resulting image set encodes both global and local shape cues, which are especially beneficial for shape completion. To fully exploit the priors, we have designed a shape fusion module for producing an initial complete shape from multi-modality input (i.e., images and point clouds), and a follow-up shape consolidation module to obtain the final complete shape by discarding unreliable points introduced by the inconsistency from diffusion priors. Extensive experimental results demonstrate our superior performance, especially in recovering fine details.
Guangshun Wei, Long Ma 0009, Chen Wang 0054, Yuanfeng Zhou, Changjian Li 0001
CVPR5
2025 Contour Makes it Stronger: Cross-Domain Cephalometric Landmark Detection Based on Contour Priors
Runnan Chen, Guangshun Wei, Shaojie Zhuang 0001, Yuanfeng Zhou
MICCAI (7)5
2025 A dynamic arrangement framework for automatic tooth alignment based on orthodontic rules
Yeying Fan, Guangshun Wei, Chuanxiang Yang, Chuanyun Fu, Wenping Wang 0001, Yuanfeng Zhou
Comput. Aided Geom. Des.7
2025 Geodesic heatmap-based segmentation-free framework for robust tooth landmark detection
Shaojie Zhuang 0001, Yeying Fan, Guangshun Wei, Yuanfeng Zhou
Comput. Graph.5
2025 Diff-TRGN: Diffusion-based tooth root generation network with multimodal clinical guidance
Chen Wang 0054, Honghao Dai, Guangshun Wei, Yuanfeng Zhou
Comput. Graph.5
2025 An improved algorithm for full-mouth lesion detection based on YOLOv8
abstract
In medical imaging detection of oral Cone Beam Computed Tomography (CBCT), there exist tiny lesions that are challenging to detect with low accuracy. The existing detection models are relatively complex. To address this, this paper presents a dual-stage YOLO detection method improved based on YOLOv8. Specifically, we first reconstruct the backbone network based on MobileNetV3 to enhance computational speed and efficiency. Second, we improve detection accuracy from three aspects: we design a composite feature fusion network to enhance the model’s feature extraction capability, addressing the issue of decreased detection accuracy for small lesions due to the loss of shallow information during the fusion process; we further combine spatial and channel information to design the C2f-SCSA module, which delves deeper into the lesion information. To tackle the problem of limited types and insufficient samples of lesions in existing CBCT images, our team collaborated with a professional dental hospital to establish a high-quality dataset, which includes 15 types of lesions and over 2000 accurately labeled oral CBCT images, providing solid data support for model training. Experimental results indicate that the improved method enhances the accuracy of the original algorithm by 3.5 percentage points, increases the recall rate by 4.7 percentage points, and raises the mean Average Precision (mAP) by 3.3 percentage points, a computational load of only 7.6 GFLOPs. This demonstrates a significant advantage in intelligent diagnosis of full-mouth lesions while improving accuracy and reducing computational load.
Xinchen Jiao, Shanshan Gao 0003, Faqiang Huang, Wenhan Dou, Yuanfeng Zhou, Caiming Zhang 0001
Graph. Model.5
2025 Collision-free path planning method for digital orthodontic treatment
abstract
The rapid evolution of digital orthodontics has highlighted a critical need for automated treatment planning systems that balance computational efficiency with clinical reliability. However, existing methods still suffer from several limitations, including excessive clinician involvement (accounting for over 35% of treatment planning time), reliance on empirically defined key frames, and limited biomechanical plausibility, particularly in cases of severe dental crowding. This paper proposes a novel collision-free optimization framework to address these issues simultaneously. Our method defines a total movement energy function evaluated over each tooth’s pose at intermediate time frames. This energy is minimized iteratively using a steepest descent strategy. A rollback mechanism is employed: if inter-tooth penetration is detected during an update, the step size is halved repeatedly until collisions are eliminated. The framework allows flexible control over the number of intermediate frames to enforce a strict constraint on per-tooth displacement, limiting it to 0.2 mm translation or 2 ° rotation every 10 to 14 days. Clinical evaluations show that the proposed algorithm can generate desirable and clinically valid tooth movement plans, even in complex cases, while significantly reducing the need for manual intervention.
Longdu Liu, Shuang-Min Chen, Lin Lu 0001, Yuanfeng Zhou, Shi-Qing Xin, Changhe Tu
Graph. Model.5
2025 Diff-OSGN: Diffusion-Based Occlusal Surface Generation Network with Geometric Constraints
abstract
Designing a functional occlusal surface for denture crowns is a complex and important task in prosthodontics. Manual design is time-consuming and heavily relies on the dentist's experience, as it requires careful consideration of occlusal function. Due to the limitations of manual design, the field has turned to data-driven methods for occlusal surface design. However, many of these methods neglect critical geometric details, such as normals and curvature, impacting the quality of the occlusal surface. In this paper, we introduce Diff-OSGN, a novel denture crown occlusal surface generation network based on a denoising diffusion model, which focuses on generating the detailed geometric structure of denture crowns. We model the occlusal surface as a geometry map based on the occlusal plane, incorporating height and normal maps rasterized from intra-oral crown scanning. Both maps represent occlusal surface geometry, and their combination further enhances these details. Considering the crucial occlusal information, we extract features from the geometry maps of adjacent and occlusal teeth, using them as conditions in the reverse diffusion process to train our network for optimal occlusal function. Additionally, we define three geometric operators and corresponding loss functions as constraints to better extract geometric features of the target occlusal surface, such as ridges and grooves, for adequate supervision. Our results demonstrate that Diff-OSGN provides quantitatively and qualitatively superior performance than competing baselines and state-of-the-art methods.
Chen Wang 0054, Guangshun Wei, James Kit Hon Tsoi, Zhiming Cui 0001, Shuyi Lu, Zhenpeng Liu, Yuanfeng Zhou
Comput. Vis. Media7
2025 Sparse Spike Feature Learning to Recognize Traceable Interictal Epileptiform Spikes
abstract
Interictal epileptiform spikes (spikes) and epileptogenic focus are strongly correlated. However, partial spikes are insensitive to epileptogenic focus, which restricts epilepsy neurosurgery. Therefore, identifying spike subtypes that are strongly associated with epileptogenic focus (traceable spikes) could facilitate their use as reliable signal sources for accurately tracing epileptogenic focus. However, the sparse firing phenomenon in the transmission of intracranial neuronal discharges leads to differences within spikes that cannot be observed visually. Therefore, neuro-electro-physiologists are unable to identify traceable spikes that could accurately locate epileptogenic focus. Herein, we propose a novel sparse spike feature learning method to recognize traceable spikes and extract discrimination information related to epileptogenic focus. First, a multilevel eigensystem feature representation was determined based on a multilevel feature representation module to express the intrinsic properties of a spike. Second, the sparse feature learning module expressed the sparse spike multi-domain context feature representation to extract sparse spike feature representations. Among them, a sparse spike encoding strategy was implemented to effectively simulate the sparse firing phenomenon for the accurate encoding of the activity of intracranial neurosources. The sensitivity of the proposed method was 97.1%, demonstrating its effectiveness and significant efficiency relative to other state-of-the-art methods.
Chenchen Cheng, Yunbo Shi, Yan Liu 0041, Yuanfeng Zhou, Ardalan Aarabi, Yakang Dai
Int. J. Neural Syst.5
2025 Monge-Ampere Regularization for Learning Arbitrary Shapes From Point Clouds
abstract
As commonly used implicit geometry representations, the signed distance function (SDF) is limited to modeling watertight shapes, while the unsigned distance function (UDF) is capable of representing various surfaces. However, its inherent theoretical shortcoming, i.e., the non-differentiability at the zero-level set, would result in sub-optimal reconstruction quality. In this paper, we propose the scaled-squared distance function (S2DF), a novel implicit surface representation for modeling arbitrary surface types. S2DF does not distinguish between inside and outside regions while effectively addressing the non-differentiability issue of UDF at the zero-level set. We demonstrate that S2DF satisfies a second-order partial differential equation of Monge-Ampere-type, allowing us to develop a learning pipeline that leverages a novel MongeAmpere regularization to directly learn S2DF from raw unoriented point clouds without supervision from ground-truth S2DF values. Extensive experiments across multiple datasets show that our method significantly outperforms state-of-the-art supervised approaches that require ground-truth surface information as supervision for training. The code will be publicly available at https://github.com/chuanxiang-yang/S2DF.
Chuanxiang Yang, Yuanfeng Zhou, Guangshun Wei, Long Ma 0009, Junhui Hou, Yuan Liu 0025, Wenping Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 RGBE-Gaze: A Large-Scale Event-Based Multimodal Dataset for High Frequency Remote Gaze Tracking
abstract
High-frequency gaze tracking demonstrates significant potential in various critical applications, such as foveatedrendering, gaze-based identity verification, and the diagnosis of mental disorders. However, existing eye-tracking systems based on CCD/CMOS cameras either provide tracking frequencies below 200 Hz or employ high-speedcameras, causing high power consumption and bulky devices. While there have been some high-speed eye-tracking datasets and methods based on event cameras, they are primarily tailored for near-eye camera scenarios. They lackthe advantages associated with remote camera scenarios, such as the absence of the need for direct contact, improved user comfort and head pose freedom. In this work, we present RGBE-Gaze, the first large-scale and multimodal dataset for remote gaze tracking in high-frequency through synchronizing RGB and event cameras. This dataset is collected from 66 participants with diverse genders and age groups. Our setup captures 3.6 million RGB images and 26.3 billion event samples. Additionally, the dataset includes 10.7 million gaze references from the Gazepoint GP3 HD eye tracker and 15,972 sparse points of gaze (PoG) ground truth obtained through manualstimuli clicks by participants. We present dataset characteristics such as head pose, gaze direction, and pupil size. Furthermore, we introduce a hybrid frame-event based gaze estimation method specifically designed for the collected dataset. Moreover, we perform extensive evaluations of different benchmarking methods under variousgaze-related factors. The evaluation results illustrate that introducing event stream as a new modality improves gazetracking frequency and demonstrates greater estimation robustness across diverse gaze-related factors.
Guangrong Zhao, Yiran Shen 0001, Zhaoxin Shen, Yuanfeng Zhou, Hongkai Wen 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Robust Hybrid Learning for Automatic Teeth Segmentation and Labeling on 3D Dental Models
abstract
Automatic teeth segmentation and labeling on dental models are basic tasks in computer-aided dentistry. Many existing works can achieve promising results in teeth segmentation, but they heavily rely on aligned input dental models, which leads to additional manual intervention. Moreover, tooth labeling is an essential task in digital dentistry for treatment planning (e.g., orthodontic), and is usually ignored in these methods. In this article, we propose an AlignNet for aligning dental models of arbitrary sizes and orientations automatically. Meanwhile, a multi-task hybrid learning network is designed that effectively plays the advantages of semantic segmentation and instance segmentation, and synergistically improves the performance of teeth point clouds segmentation and labeling. Particularly, for the teeth-gingival boundaries with large segmentation errors, we utilize the filtered curvature information as a constrained feature to detect the weak boundary more accurately. At last, we propose a DiffLoss and postprocessing step based on the dental arch to address the teeth classification problem. Through extensive evaluations of oral scanning models, our method is robust to handle dental model point clouds with arbitrary size and orientation, and outperforms state-of-the-art teeth segmentation and labeling methods, demonstrating its full automation and robustness in clinical practice.
Shaojie Zhuang 0001, Guangshun Wei, Zhiming Cui 0001, Yuanfeng Zhou
IEEE Trans. Multim.4
2025 NeuVAS: Neural Implicit Surfaces for Variational Shape Modeling
abstract
Neural implicit shape representation has drawn significant attention in recent years due to its smoothness, differentiability, and topological flexibility. However, directly modeling the shape of a neural implicit surface, especially as the zero-level set of a neural signed distance function (SDF), with sparse geometric control is still a challenging task. Sparse input shape control typically includes 3D curve networks or, more generally, 3D curve sketches, which are unstructured and cannot be connected to form a curve network, and therefore more difficult to deal with. While 3D curve networks or curve sketches provide intuitive shape control, their sparsity and varied topology pose challenges in generating high-quality surfaces to meet such curve constraints. In this paper, we propose NeuVAS, a variational approach to shape modeling using neural implicit surfaces constrained under sparse input shape control, including unstructured 3D curve sketches as well as connected 3D curve networks. Specifically, we introduce a smoothness term based on a functional of surface curvatures to minimize shape variation of the zero-level set surface of a neural SDF. We also develop a new technique to faithfully model G 0 sharp feature curves as specified in the input curve sketches. Comprehensive comparisons with the state-of-the-art methods demonstrate the significant advantages of our method.
Qiujie Dong, Fangtian Liang, Hao Pan 0001, Lei Yang 0048, Congyi Zhang 0001, Guying Lin, Caiming Zhang 0001, Yuanfeng Zhou, Changhe Tu, Shi-Qing Xin, Alla Sheffer, Xin Li 0003, Wenping Wang 0001
ACM Trans. Graph.9
2025 D-FRAME: Direction-Field-Based Wireframe Extraction for Complex CAD Models
abstract
Extracting wireframes from CAD models represented by point cloud remains a significant challenge in computer graphics. This difficulty arises from two main factors: first, imperfections in the point cloud data, such as lack of orientation, noise, and sparsity; and second, the inherent complexity of geometric shapes, which often feature a high density of sharp edges in close proximity. In this paper, we propose D-FRAME, a multi-stage wireframe extraction framework that incorporates a novel direction field to improve edge detection quality and connectivity, a refinement strategy to address sparse or noisy edge points, and a final coarse-to-fine connection module to extract a robust wireframe. The direction field not only facilitates connectivity but also enhances the precision of extracted edges by mitigating the impact of misclassified points. By combining the Restricted Voronoi Diagram (RVD) with the extracted wireframes and the original point cloud, our approach also achieves highly faithful reconstruction of CAD model. Experiments conducted on synthetic and real-world scanned CAD datasets demonstrate that D-FRAME effectively manages noise, sparsity, and complex geometries, yielding high-fidelity wireframes.
Honghao Dai, Guangshun Wei, Long Ma 0009, Yuanfeng Zhou, Ying He 0001
IEEE Trans. Vis. Comput. Graph.6
2025 A Potential Field Method for Tooth Motion Planning in Orthodontic Treatment
abstract
Invisible orthodontics, commonly known as clear alignment treatment, offers a more comfortable and aesthetically pleasing alternative in orthodontic care, attracting considerable attention in the dental community in recent years. It replaces conventional metal braces with a series of removable, and transparent aligners. Each aligner is crafted to facilitate a gradual adjustment of the teeth, ensuring progressive stages of dental correction. This necessitates the design for teeth motion. Here we present an automatic method and a system for generating collision-free teeth motion planning while avoiding gaps between adjacent teeth, which is unacceptable in clinical practice. To tackle this task, we formulate it as a constrained optimization problem and utilize the interior point method for its solution. We also developed an interactive system that enables dentists to easily visualize and edit the paths. Our method significantly speeds up the clear aligner planning process, creating the desired motion paths for a full set of teeth in under five minutes-a task that typically requires several hours of manual work. Our experiments and user studies confirm the effectiveness of this method in planning teeth movement, showcasing its potential to streamline orthodontic procedures.
Yuexin Ma, Lei Yang 0048, Congyi Zhang 0001, Guangshun Wei, Runnan Chen, Min Gu 0003, Jia Pan 0001, Zhengbao Yang, Taku Komura, Shi-Qing Xin, Yuanfeng Zhou, Changhe Tu, Wenping Wang 0001
IEEE Trans. Vis. Comput. Graph.12
2025 Computing Smooth and Integrable Cross Fields via Iterative Singularity Adjustment
abstract
We propose a new method for computing smooth and integrable cross fields on 2D and 3D surfaces. our approach first computes smooth cross fields by minimizing the Dirichlet energy. Unlike existing optimization-based methods, our technique determines the singularity configuration-i.e., the number, locations, and indices of singularities-by iteratively adjusting them. Singularities can move, merge and split, akin to the behavior of like charges repelling and unlike charges attracting. Once all singularities stop moving, we obtain a cross field with (locally) the lowest Dirichlet energy. In simply connected domains, this cross field is guaranteed to be integrable. However, this property does not hold in multiply connected domains. To make a smooth cross field integrable, we construct a vector field $\bf c$c that characterizes the deviation of the cross field from a curl-free field. We then optimize the locations of singularities by moving them along the field lines of $\bf c$c. Our method is fundamentally different from existing integer programming-based approaches, as it avoids combinatorial optimization. It is fully automatic and includes a parameter to control the number of singularities. Our method is well suited for smooth models where exact boundary alignment and sparse hard directional constraints are desired, and can guide seamless conformal parameterization and T-junction-free quadrangulation.
Long Ma 0009, Ying He 0001, Jianmin Zheng, Yuanfeng Zhou, Shi-Qing Xin, Caiming Zhang 0001, Wenping Wang 0001
IEEE Trans. Vis. Comput. Graph.4
2025 A Rule-Based Optimization Method for Tooth Alignment
abstract
While tooth alignment is crucial for digital dentistry, especially in orthodontic treatment, existing computer-aided methods mainly focus on the 3D dental crown but overlook the entire teeth, which is essential for applications in orthodontics. Besides, clinical orthodontic rules are not fully considered in these methods, i.e., there should be no collisions and gaps between teeth, the upper jaw and lower jaw should have correct occlusion relationships, the teeth should comply with a reasonable dental arch curve, etc. To generate optimal tooth alignment results, we propose a rule-based optimization method for solving the tooth alignment problem that takes into consideration the clinical rules functionally and aesthetically. We optimize rule-driven objective functions by adjusting the 6-DoF transformations of each tooth. Besides, our optimization formulation supports customization for different clinical scenarios by specifying the various energy terms. Extensive experiments, ablation studies, and user studies have been conducted to validate the effectiveness of our method. Quantitative and qualitative comparisons demonstrate that our method generates better tooth alignments than previous methods.
Yuhan Ping, Guodong Wei, Guangshun Wei, Congyi Zhang 0001, Noha A. SAID, Jia Pan 0001, Shi-Qing Xin, Yuanfeng Zhou, Changhe Tu, Min Gu 0003, Wenping Wang 0001
IEEE Trans. Vis. Comput. Graph.8
2025 Design and Optimization of Self-Supporting Surfaces With Arch Beams
abstract
The article presents a new method for constructing self-supporting surfaces using arch beams that are designed to convert their thrust into supporting force, thereby eliminating shear stress and bending moments. Our method allows for the placement of the arch beams on the boundary or within a surface and partitions the surface into multiple self-supporting parts. The use of arch beams enhances stability and durability, adds aesthetic appeal, and allows for greater flexibility in the design process. We develop an iterative algorithm for designing self-supporting surfaces with arch beams that enables the user to control the shape of the beams and surface through intuitive parameters and specify the desired location of the arch beams. We verify the physical stability of the structure using finite element analysis. Experimental results show that our method can produce visually pleasing self-supporting surfaces that satisfy the equilibrium equation with high accuracy.
Guangshun Wei, Long Ma 0009, Yuanfeng Zhou, Chen Wang 0054, Jianmin Zheng, Ying He 0001
IEEE Trans. Vis. Comput. Graph.3
2025 AirtypeLogger: How Short Keystrokes in Virtual Space Can Expose Your Semantic Input to Nearby Cameras
abstract
Considering the issue of privacy leakage and motivating more sophisticated protection methods for air-typing with XR devices, in this paper, we propose AirtypeLogger, a new approach towards practical video-based attacks on the air-typing activities of XR users in virtual space. Different from the existing approaches, AirtypeLogger considers a scenario in which the users are typing a short text fragment with semantic meaning occasionally under the spy of video cameras. It detects and localizes the air-typing events in video streams and proposes the spatial-temporal representation to encode the keystrokes' relative positions and temporal order. Then, high-precision inference can be achieved by applying a Transformer-based network to the spatial and temporal encodings of the keystroke sequences. Finally, according to our extensive real-world experiments, AirtypeLogger can achieve a Character Error Rate (CER) of less than 0.1 as long as 7 air-typing events are observed, which is impossible for previous approaches that require long-term observation of the typing activities online before launching inference attacks. The implementation details and source codes can be found at https://github.com/ztysdu/AirtypeLogger.
Tongyu Zhang, Yiran Shen 0001, Yuanfeng Zhou
IEEE Trans. Vis. Comput. Graph.5
2024 Collaborative Tooth Motion Diffusion Model in Digital Orthodontics
abstract
Tooth motion generation is an essential task in digital orthodontic treatment for precise and quick dental healthcare, which aims to generate the whole intermediate tooth motion process given the initial pathological and target ideal tooth alignments. Most prior works for multi-agent motion planning problems usually result in complex solutions. Moreover, the occlusal relationship between upper and lower teeth is often overlooked. In this paper, we propose a collaborative tooth motion diffusion model. The critical insight is to remodel the problem as a diffusion process. In this sense, we model the whole tooth motion distribution with a diffusion model and transform the planning problem into a sampling process from this distribution. We design a tooth latent representation to provide accurate conditional guides consisting of two key components: the tooth frame represents the position and posture, and the tooth latent shape code represents the geometric morphology. Subsequently, we present a collaborative diffusion model to learn the multi-tooth motion distribution based on inter-tooth and occlusal constraints, which are implemented by graph structure and new loss functions, respectively. Extensive qualitative and quantitative experiments demonstrate the superiority of our framework in the application of orthodontics compared with state-of-the-art methods.
Yeying Fan, Guangshun Wei, Chen Wang 0054, Shaojie Zhuang 0001, Wenping Wang 0001, Yuanfeng Zhou
AAAI6
2024 Utilizing Genetic Feature Augmentation and CAM Consistency for Enhanced Classification of Imbalanced Skin Lesion Data
abstract
Automated classification of skin lesions has been proven to improve the diagnostic accuracy of dermoscopic images significantly. Despite numerous achievements in this field, accurate classification remains challenging due to the diversity of skin cancer lesions and the class imbalance within datasets. Traditional methods typically augment imbalanced datasets by increasing perturbations to input data, but the improvements are often limited. Therefore, this paper proposes a method that utilizes genetic feature augmentation and class activation mapping (CAM) consistency for enhanced classification of imbalanced skin lesion data. The proposed genetic feature module incorporates the concept of genetic algorithms into the network model. The model can learn more diverse expressions by selectively retaining highly representative features while cross-replicating data from underrepresented classes. Additionally, we introduce a global lesion localization module based on CAM to enhance the learning of discriminative features among classes. This module optimizes the multiple CAM distances generated from the same dermoscopic image, thereby improving the differentiation of inter-class features. To validate the effectiveness of this method, extensive experiments were conducted on the ISIC-2017 and ISIC-2018 datasets. The experimental results demonstrate that this method performs exceptionally well on multi-class data, significantly improving classification accuracy and stability.
Shuyi Lu, Yuanfeng Zhou
BIBM3
2024 Low-Quality Ultrasound Image Enhancement Using Spatially Variable Neural Implicit Network
abstract
Ultrasound images are widely used because they are easy to use and non-invasive. However, differences in acquisition equipment and frequency settings can result in varying image quality. To address this issue, we propose an ultrasound image enhancement method based on spatially variable neural implicit networks. Our approach consists of several steps: First, the encoder of an arbitrary-scale super-resolution network processes low-quality ultrasound images to generate a feature matrix. Concurrently, the parameter generation network decomposes the low-quality image into a parameter matrix. The pixel prediction module of the arbitrary-scale super-resolution network then uses the parameter matrix to map encoded features to pixel values. To further align the predicted pixel values with the distribution of high-quality ultrasound images and reduce interference from artifacts, we design a spectrum discriminator network based on the fast Fourier transform. This network is highly effective at capturing high-quality information and minimizing the impact of redundant information. In both qualitative and quantitative comparisons with current state-of-the-art image enhancement methods and arbitrary magnification techniques, our approach demonstrates superior performance.
Shanshan Gao 0003, Yuanfeng Zhou
BIBM5
2024 Automated placement of dental attachments based on orthodontic pathways
Yiheng Lv, Guangshun Wei, Yeying Fan, Long Ma 0009, Yuanfeng Zhou
Comput. Aided Geom. Des.6
2024 Construction of the ellipse with maximum area inscribed in an arbitrary convex quadrilateral
Long Ma 0009, Yuanfeng Zhou
Comput. Aided Geom. Des.2
2024 High-precision teeth reconstruction based on automatic multimodal fusion with CBCT and IOS
Long Ma 0009, Minfeng Xu, Guangshun Wei, Shaojie Zhuang 0001, Yuanfeng Zhou
Comput. Aided Geom. Des.6
2024 Decoupled and boosted learning for skeleton-based dynamic hand gesture recognition
Yangke Li, Guangshun Wei, Christian Desrosiers, Yuanfeng Zhou
Pattern Recognit.4
2024 Aggregating Global and Local Representations via Hybrid Transformer for Video Deraining
abstract
Although video deraining technology has achieved great success in recent years, extracting spatiotemporal feature representations across the domains of spatial and temporal in successive frames, then performing spatial and temporal modeling, and restoring high-quality deraining videos with rich details are still challenging tasks. In this paper, we use the hybrid Transformer for the first attempt in video rain removal tasks, and propose a novel video deraining network based on hybrid transformer (VDN-HT) to aggregate global and local representations to accomplish video deraining. In the feature extraction process, we propose to use a U-shaped structure based on serial Transformer blocks to extract shallow local features, deep global features and global dependencies, and then adaptively aggregate them to obtain rainy video features with rain streaks of different directions and densities. In order to better model spatiotemporal relationships, the VDN-HT uses the Transformer’s long-range and relational modeling abilities to obtain the features of spatial and the correlations of temporal between continuous video frames to achieve multi-frame alignment. For ensuring the global-local consistency of the reconstructed frames, we design a global-local reconstruction module composed of Transformer and convolutional neural network (CNN) in parallel to aggregate global and local information to better reconstruct each frame. In addition, the proposed gating-based refinement module and color loss effectively retain the details and color information after removing rain streaks. Extensive experiments on NTURain, RainSynLight25 and RainSynHeavy25 datasets have shown that the VDN-HT can handle many types of rainy videos and perform better than previous methods.
Deqian Mao, Shanshan Gao 0003, Honghao Dai, Yunfeng Zhang 0001, Yuanfeng Zhou
IEEE Trans. Circuits Syst. Video Technol.6
2024 Semi-Supervised Medical Image Segmentation Using Cross-Style Consistency With Shape-Aware and Local Context Constraints
abstract
Despite the remarkable progress in semi-supervised medical image segmentation methods based on deep learning, their application to real-life clinical scenarios still faces considerable challenges. For example, insufficient labeled data often makes it difficult for networks to capture the complexity and variability of the anatomical regions to be segmented. To address these problems, we design a new semi-supervised segmentation framework that aspires to produce anatomically plausible predictions. Our framework comprises two parallel networks: shape-agnostic and shape-aware networks. These networks learn from each other, enabling effective utilization of unlabeled data. Our shape-aware network implicitly introduces shape guidance to capture shape fine-grained information. Meanwhile, shape-agnostic networks employ uncertainty estimation to further obtain reliable pseudo-labels for the counterpart. We also employ a cross-style consistency strategy to enhance the network's utilization of unlabeled data. It enriches the dataset to prevent overfitting and further eases the coupling of the two networks that learn from each other. Our proposed architecture also incorporates a novel loss term that facilitates the learning of the local context of segmentation by the network, thereby enhancing the overall accuracy of prediction. Experiments on three different datasets of medical images show that our method outperforms many excellent semi-supervised segmentation methods and outperforms them in perceiving shape. The code can be seen at https://github.com/igip-liu/SLC-Net.
Jinhua Liu 0003, Christian Desrosiers, Dexin Yu, Yuanfeng Zhou
IEEE Trans. Medical Imaging4
2024 An Adaptive Sample Assignment Network for Tiny Object Detection
abstract
Tiny objects often have a small proportion of pixels in the image, leading to significant differences in the number of positive and negative samples and the lack of feature information. Accurately determining the position and category of tiny objects remains a huge challenge for object detection research. Therefore, we design an Adaptive Sample Assignment Strategy(ASAS) and tiny object focusing enhancement module to solve the above two problems. Specifically, starting from the study of positive and negative sample selection and balance strategies for tiny objects, we construct a lightweight Object Existence Probability Determination Network (OEPD/Net) to focus on the areas where tiny objects exist, and achieve adaptive assignment and balance of samples. A top/down, layer by layer focusing enhancement module is designed to effectively enhance the propagation ability of high/level semantic information for tiny objects. The above two solutions have excellent generalization and migration capabilities and can be applied to any stage and two-stage object detection network, effectively enhancing TOD performance. Finally, this article provides a performance analysis of detection performance the detection network based on the OEPD/Net output results, and demonstrates the effectiveness of the proposed OEPD-Net and focusing enhancement module through extensive experiments on a public dataset.
Honghao Dai, Shanshan Gao 0003, Deqian Mao, Chenhao Zhang 0001, Yuanfeng Zhou
IEEE Trans. Multim.6
2024 Deep Plug-and-Play Non-Iterative Cluster for 3D Global Feature Extraction
abstract
Efficient and accurate point cloud feature extraction is crucial for critical tasks such as 3D recognition and semantic segmentation. However, existing global feature extraction methods for 3D data often require designing different models for different input types (point clouds, voxels, and maps). This article proposes an efficient plug-and-play non-iterative clustering method (NICM) to establish a unified point cloud global feature extraction paradigm suitable for any input type to solve the above problems. The core idea of the NICM is to construct the connection between a single point and other global points based only on the cosine similarity between center points to achieve global feature extraction, which has linear complexity characteristics and can be combined with any existing feature extraction model. Additionally, to better integrate the features extracted by NICM and the original model, this article designs an adaptive feature fusion module is designed based on the gate unit, which retains similar features and effectively fuses dissimilar features based on their importance to downstream tasks. We have applied our method to downstream tasks such as point cloud recognition, part segmentation and scene segmentation. Sufficient experiments have proven that our method can provide comprehensive and robust features for the original model, and effectively improve the performance of downstream tasks.
Shanshan Gao 0003, Deqian Mao, Shouwen Song, Lei Li 0008, Yuanfeng Zhou
ACM Trans. Multim. Comput. Commun. Appl.6
2024 VibHead: An Authentication Scheme for Smart Headsets through Vibration
abstract
Recent years have witnessed the fast penetration of Virtual Reality (VR) and Augmented Reality (AR) systems into our daily life, the security and privacy issues of the VR/AR applications have been attracting considerable attention. Most VR/AR systems adopt head-mounted devices (i.e., smart headsets) to interact with users and the devices usually store the users’ private data. Hence, authentication schemes are desired for the head-mounted devices. Traditional knowledge-based authentication schemes for general personal devices have been proved vulnerable to shoulder-surfing attacks, especially considering the headsets may block the sight of the users. Although the robustness of the knowledge-based authentication can be improved by designing complicated secret codes in virtual space, this approach induces a compromise of usability. Another choice is to leverage the users’ biometrics; however, it either relies on highly advanced equipments which may not always be available in commercial headsets or introduce heavy cognitive load to users. In this paper, we propose a vibration-based authentication scheme, VibHead, for smart headsets. Since the propagation of vibration signals through human heads presents unique patterns for different individuals, VibHead employs a CNN-based model to classify registered legitimate users based the features extracted from the vibration signals. We also design a two-step authentication scheme where the above user classifiers are utilized to distinguish the legitimate user from illegitimate ones. We implement VibHead on a Microsoft HoloLens equipped with a linear motor and an IMU sensor which are commonly used in off-the-shelf personal smart devices. According to the results of our extensive experiments, with short vibration signals (≤ 1s ), VibHead has an outstanding authentication accuracy; both FAR and FRR are around 5%.
Feng Li 0002, Huan Yang 0001, Dongxiao Yu, Yuanfeng Zhou, Yiran Shen 0001
ACM Trans. Sens. Networks5
2024 Tooth Alignment Network Based on Landmark Constraints and Hierarchical Graph Structure
abstract
Automatic tooth alignment target prediction is vital in shortening the planning time of orthodontic treatments and aligner designs. Generally, the quality of alignment targets greatly depends on the experience and ability of dentists and has enormous subjective factors. Therefore, many knowledge-driven alignment prediction methods have been proposed to help inexperienced dentists. Unfortunately, existing methods tend to directly regress tooth motion, which lacks clinical interpretability. Tooth anatomical landmarks play a critical role in orthodontics because they are effective in aiding the assessment of whether teeth are in close arrangement and normal occlusion. Thus, we consider anatomical landmark constraints to improve tooth alignment results. In this article, we present a novel tooth alignment neural network for alignment target predictions based on tooth landmark constraints and a hierarchical graph structure. We detect the landmarks of each tooth first and then construct a hierarchical graph of jaw-tooth-landmark to characterize the relationship between teeth and landmarks. Then, we define the landmark constraints to guide the network to learn the normal occlusion and predict the rigid transformation of each tooth during alignment. Our method achieves better results with the architecture built for tooth data and landmark constraints and has better explainability than previous methods with regard to clinical tooth alignments.
Chen Wang 0054, Guangshun Wei, Guodong Wei, Wenping Wang 0001, Yuanfeng Zhou
IEEE Trans. Vis. Comput. Graph.5
2024 iPUNet: Iterative Cross Field Guided Point Cloud Upsampling
abstract
Point clouds acquired by 3D scanning devices are often sparse, noisy, and non-uniform, causing a loss of geometric features. To facilitate the usability of point clouds in downstream applications, given such input, we present a learning-based point upsampling method, i.e., iPUNet, which generates dense and uniform points at arbitrary ratios and better captures sharp features. To generate feature-aware points, we introduce cross fields that are aligned to sharp geometric features by self-supervision to guide point generation. Given cross field defined frames, we enable arbitrary ratio upsampling by learning at each input point a local parameterized surface. The learned surface consumes the neighboring points and 2D tangent plane coordinates as input, and maps onto a continuous surface in 3D where arbitrary ratios of output points can be sampled. To solve the non-uniformity of input points, on top of the cross field guided upsampling, we further introduce an iterative strategy that refines the point distribution by moving sparse points onto the desired continuous 3D surface in each iteration. Within only a few iterations, the sparse points are evenly distributed and their corresponding dense samples are more uniform and better capture geometric features. Through extensive evaluations on diverse scans of objects and scenes, we demonstrate that iPUNet is robust to handle noisy and non-uniformly distributed inputs, and outperforms state-of-the-art point cloud upsampling methods.
Guangshun Wei, Hao Pan 0001, Shaojie Zhuang 0001, Yuanfeng Zhou, Changjian Li 0001
IEEE Trans. Vis. Comput. Graph.4
2024 Swift-Eye: Towards Anti-blink Pupil Tracking for Precise and Robust High-Frequency Near-Eye Movement Analysis with Event Cameras
abstract
Eye tracking has shown great promise in many scientific fields and daily applications, ranging from the early detection of mental health disorders to foveated rendering in virtual reality (VR). These applications all call for a robust system for high-frequency near-eye movement sensing and analysis in high precision, which cannot be guaranteed by the existing eye tracking solutions with CCD/CMOS cameras. To bridge the gap, in this paper, we propose Swift-Eye, an offline precise and robust pupil estimation and tracking framework to support high-frequency near-eye movement analysis, especially when the pupil region is partially occluded. Swift-Eye is built upon the emerging event cameras to capture the high-speed movement of eyes in high temporal resolution. Then, a series of bespoke components are designed to generate high-quality near-eye movement video at a high frame rate over kilohertz and deal with the occlusion over the pupil caused by involuntary eye blinks. According to our extensive evaluations on EV-Eye, a large-scale public dataset for eye tracking using event cameras, Swift-Eye shows high robustness against significant occlusion. It can improve the IoU and F1-score of the pupil estimation by 20% and 12.5% respectively, compared with the second-best competing approach, when over 80% of the pupil region is occluded by the eyelid. Lastly, it provides continuous and smooth traces of pupils in extremely high temporal resolution and can support high-frequency eye movement analysis and a number of potential applications, such as mental health diagnosis, behaviour-brain association, etc. The implementation details and source codes can be found at https://github.com/ztysdu/Swift-Eye.
Tongyu Zhang, Yiran Shen 0001, Guangrong Zhao, Lin Wang 0025, Xiaoming Chen 0006, Lu Bai 0004, Yuanfeng Zhou
IEEE Trans. Vis. Comput. Graph.7
2023 Interactive Segmentation for Pathological Images with Similarity-based Propagation
abstract
The auxiliary diagnosis based on pathological images often requires detecting exact nuclear information. In this paper, we propose an iteratively-refined interactive segmentation network named PSINet that allows users to guide the segmentation process of the model by drawing scribbles. PSINet can learn long-range dependencies among different cell nuclei, allowing it to correct other nuclei without direct feedback when the user only corrects the segmentation results on a few nuclei. Experimental results show that the proposed network outperforms state-of-the-art iteratively-refined interactive segmentation networks.
Jinhua Liu 0003, Guangshun Wei, Yuanfeng Zhou
BIBM5
2023 Deep Residual Fourier and Self-Attention for Arbitrary Scale MRI Super-Resolution
abstract
The excellent inherent contrast between biological tissues afforded by MR imaging is one of the foremost characteristics of this technique, but it depends on the scan time and hardware devices. Recent studies have shown significant progress in deep learning-based single-image super-resolution (SISR) algorithms based on MR images. Many researchers have employed implicit functions for super-resolution tasks, achieving arbitrary resolution upsampling. Nevertheless, challenges like loss of generated image texture and weak high-frequency features persist. We propose an arbitrary multiple super-resolution model based on deep residual Fourier transform and a selfattention mechanism to tackle these issues. We first use the deep Fourier residual block to construct the high and low-frequency difference of the image, compensating for the spectral bias of the multiple perceptron (MLP) network. Building upon the substantial similarity in medical image tissue structures, we add a vertical and horizontal self-attention mechanism to capture the internal correlation of features in the vertical and horizontal directions. Finally, we learn a continuous functional representation to execute super-resolution tasks at any scale. Experiment results show the effectiveness of our method on the T2 sequence of public dataset SIMON and T1, T2, and Proton Density (PD) weighted scan sequences of public dataset IXI. We conduct qualitative and quantitative comparisons of the three contrasts, illustrating the superiority of our approach.
Shanshan Gao 0003, Xingwei Hao, Yuanfeng Zhou
BIBM5
2023 Metaballs-Based Real-Time Elastic Object Simulation via Projective Dynamics
Shanshan Gao 0003, Yuanfeng Zhou
CAD/Graphics5
2023 Hybrid Optimization-based Cutting Simulation for Soft Objects
Long Ma 0009, Minfeng Xu, Yuanfeng Zhou
Comput. Aided Des.5
2023 A Region-growing GradNormal Algorithm for Geometrically and Topologically Accurate Mesh Extraction
Chen Zong, Jinhui Zhao, Shuang-Min Chen, Shi-Qing Xin, Yuanfeng Zhou, Changhe Tu, Wenping Wang 0001
Comput. Aided Des.6
2023 A new method for researching and constructing spherical bicentric polygons based on geometric mapping
Long Ma 0009, Yuanfeng Zhou
Comput. Aided Geom. Des.3
2023 Weak-Boundary Sensitive Superpixel Segmentation Based on Local Adaptive Distance
abstract
Superpixel segmentation provides a way to capture object boundaries unsupervised and has benefited many compute vision applications. However, under-segmentation for weak boundaries and poor compatibility with image feature representations often limit its wide application. In this paper, we propose a new weak-boundary sensitive superpixel generation method and provide an all-in-one solution for images with different feature representations. We first design a local adaptive distance (LAD) to be more sensitive to feature changes in low-contrast regions. LAD leverages image local standard deviation as region contrast clues. It adaptively increases the feature distances in low-contrast regions to avoid feature space distances of weak boundaries being inundated by regularity constraints. LAD is scale-invariant that can be compatible with high bit-depth and multi-feature images. Then, based on LAD, we introduce a novel morphological contour evolution model to generate superpixels iteratively. Leveraging morphological dilation of superpixel shapes, the new model is more conducive to the boundary detection of irregular or slender objects. Extensive experiments demonstrate that our method favorably outperforms state-of-the-art methods, especially regarding the under-segmentation error and segmentation accuracy.
Limin Sun 0003, Dongyang Ma, Yuanfeng Zhou
IEEE Trans. Circuits Syst. Video Technol.4
2023 CAT: Constrained Adversarial Training for Anatomically-Plausible Semi-Supervised Segmentation
abstract
Deep learning models for semi-supervised medical image segmentation have achieved unprecedented performance for a wide range of tasks. Despite their high accuracy, these models may however yield predictions that are considered anatomically impossible by clinicians. Moreover, incorporating complex anatomical constraints into standard deep learning frameworks remains challenging due to their non-differentiable nature. To address these limitations, we propose a Constrained Adversarial Training (CAT) method that learns how to produce anatomically plausible segmentations. Unlike approaches focusing solely on accuracy measures like Dice, our method considers complex anatomical constraints like connectivity, convexity, and symmetry which cannot be easily modeled in a loss function. The problem of non-differentiable constraints is solved using a Reinforce algorithm which enables to obtain a gradient for violated constraints. To generate constraint-violating examples on the fly, and thereby obtain useful gradients, our method adopts an adversarial training strategy which modifies training images to maximize the constraint loss, and then updates the network to be robust to these adversarial examples. The proposed method offers a generic and efficient way to add complex segmentation constraints on top of any segmentation network. Experiments on synthetic data and four clinically-relevant datasets demonstrate the effectiveness of our method in terms of segmentation accuracy and anatomical plausibility.
Ping Wang 0016, Jizong Peng, Marco Pedersoli, Yuanfeng Zhou, Caiming Zhang 0001, Christian Desrosiers
IEEE Trans. Medical Imaging4
2023 Shape-Aware Joint Distribution Alignment for Cross-Domain Image Segmentation
abstract
We present an unsupervised domain adaptation method for image segmentation which aligns high-order statistics, computed for the source and target domains, encoding domain-invariant spatial relationships between segmentation classes. Our method first estimates the joint distribution of predictions for pairs of pixels whose relative position corresponds to a given spatial displacement. Domain adaptation is then achieved by aligning the joint distributions of source and target images, computed for a set of displacements. Two enhancements of this method are proposed. The first one uses an efficient multi-scale strategy that enables capturing long-range relationships in the statistics. The second one extends the joint distribution alignment loss to features in intermediate layers of the network by computing their cross-correlation. We test our method on the task of unpaired multi-modal cardiac segmentation using the Multi-Modality Whole Heart Segmentation Challenge dataset and prostate segmentation task where images from two datasets are taken as data in different domains. Our results show the advantages of our method compared to recent approaches for cross-domain image segmentation. Code is available at https://github.com/WangPing521/Domain_adaptation_shape_prior.
Ping Wang 0016, Jizong Peng, Marco Pedersoli, Yuanfeng Zhou, Caiming Zhang 0001, Christian Desrosiers
IEEE Trans. Medical Imaging4
2023 Collaborative Multi-Metadata Fusion to Improve the Classification of Lumbar Disc Herniation
abstract
Computed tomography (CT) images are the most commonly used radiographic imaging modality for detecting and diagnosing lumbar diseases. Despite many outstanding advances, computer-aided diagnosis (CAD) of lumbar disc disease remains challenging due to the complexity of pathological abnormalities and poor discrimination between different lesions. Therefore, we propose a Collaborative Multi-Metadata Fusion classification network (CMMF-Net) to address these challenges. The network consists of a feature selection model and a classification model. We propose a novel Multi-scale Feature Fusion (MFF) module that can improve the edge learning ability of the network region of interest (ROI) by fusing features of different scales and dimensions. We also propose a new loss function to improve the convergence of the network to the internal and external edges of the intervertebral disc. Subsequently, we use the ROI bounding box from the feature selection model to crop the original image and calculate the distance features matrix. We then concatenate the cropped CT images, multiscale fusion features, and distance feature matrices and input them into the classification network. Next, the model outputs the classification results and the class activation map (CAM). Finally, the CAM of the original image size is returned to the feature selection network during the upsampling process to achieve collaborative model training. Extensive experiments demonstrate the effectiveness of our method. The model achieved 91.32% accuracy in the lumbar spine disease classification task. In the labelled lumbar disc segmentation task, the Dice coefficient reaches 94.39%. The classification accuracy in the Lung Image Database Consortium and Image Database Resource Initiative (LIDC-IDRI) reaches 91.82%.
Shuyi Lu, Jinhua Liu 0003, Yuanfeng Zhou
IEEE Trans. Medical Imaging4
2022 Semi-supervised Medical Image Segmentation Using Cross-Model Pseudo-Supervision with Shape Awareness and Local Context Constraints
Jinhua Liu 0003, Christian Desrosiers, Yuanfeng Zhou
MICCAI (8)3
2022 Dense representative tooth landmark/axis detection network on 3D model
Guangshun Wei, Zhiming Cui 0001, Lei Yang 0048, Yuanfeng Zhou, Pradeep Singh 0003, Min Gu 0003, Wenping Wang 0001
Comput. Aided Geom. Des.5
2022 TAD-Net: tooth axis detection network based on rotation transformation encoding
Yeying Fan, Guangshun Wei, Zhiming Cui 0001, Yuanfeng Zhou, Wenping Wang 0001
Graph. Model.5
2022 Constructing self-supporting surfaces with planar quadrilateral elements
abstract
We present a simple yet effective method for constructing 3D self-supporting surfaces with planar quadrilateral (PQ) elements. Starting with a triangular discretization of a self-supporting surface, we first compute the principal curvatures and directions of each triangular face using a new discrete differential geometry approach, yielding more accurate results than existing methods. Then, we smooth the principal direction field to reduce the number of singularities. Next, we partition all faces into two groups in terms of principal curvature difference. For each face with small curvature difference, we compute a stretch matrix that turns the principal directions into a pair of conjugate directions. For the remaining triangular faces, we simply keep their smoothed principal directions. Finally, applying a mixed-integer programming solver to the mixed principal and conjugate direction field, we obtain a planar quadrilateral mesh. Experimental results show that our method is computationally efficient and can yield high-quality PQ meshes that well approximate the geometry of the input surfaces and maintain their self-supporting properties.
Long Ma 0009, Sidan Yao, Jianmin Zheng, Yang Liu 0014, Yuanfeng Zhou, Shi-Qing Xin, Ying He 0001
Comput. Vis. Media5
2022 Multiview Feature Fusion Representation for Interictal Epileptiform Spikes Detection
abstract
Interictal epileptiform spikes (IES) of scalp electroencephalogram (EEG) signals have a strong relation with the epileptogenic region. Since IES are highly unlikely to be detected in scalp EEG signals, the primary diagnosis depends heavily on the visual evaluation of IES. However, visual inspection of EEG signals, the standard IES detection procedure is time-consuming, highly subjective, and error-prone. Furthermore, the highly complex, nonlinear, and nonstationary characteristics of EEG signals lead to the incomplete representation of EEG signals in existing computer-aided methods and consequently unsatisfactory detection performance. Therefore, a novel multiview feature fusion representation (MVFFR) method was developed and combined with a robustness classifier to detect EEG signals with/without IES. MVFFR comprises two steps: First, temporal, frequency, temporal-frequency, spatial, and nonlinear domain features are transformed by the IES to express the latent information effectively. Second, the unsupervised infinite feature-selection method determines the most distinct feature fusion representations. Experimental results using a balanced dataset of six patients showed that MVFFR achieved the optimal detection performance (accuracy: 89.27%, sensitivity: 89.01%, specificity: 89.54%, and precision: 89.82%) compared with other feature ranking methods, and the MVFFR-related method were complementary and indispensable. Additionally, in an independent test, MVFFR maintained excellent generalization capacity with a false detection rate per minute of 0.15 on the unbalanced dataset of one patient.
Chenchen Cheng, Yuanfeng Zhou, Yan Liu 0041, Gao Fei, Liling Yang, Yakang Dai
Int. J. Neural Syst.2
2022 Tooth instance segmentation based on capturing dependencies and receptive field adjustment in cone beam computed tomography
abstract
Abstract Automatic and accurate instance segmentation of teeth can provide important support for computer‐aided orthodontic work. Traditional methods for tooth segmentation studies often ignore the rich structural features of teeth. Capturing the complete and accurate geometry as well as morphological details of a single tooth remains a challenge for current tooth segmentation studies. In this article, a new tooth segmentation deeplearning network based on capturing dependencies and receptive field adjustment in cone beam computed tomography (CBCT) is proposed to achieve automatic and accurate instance segmentation of dental CBCT data. The method acquires coarse‐level features of tooth and accurate tooth centroids in the first stage, and acquires the instance information and spatial position localization of the tooth. The encoding process in the second stage of the network introduces a guidance module for obtaining tooth geometry information based on a 3D self‐attention mechanism to capture dependencies in CBCT. The proposed tooth feature integration module is based on multiscale fusion of dilated convolutions to capture tooth detailed information at multiple scales, and the network receptive field was adjusted. Extensive evaluation, ablation, and comparison experiments demonstrate that our method exhibits state‐of‐the‐art segmentation performance and accurate instance segmentation results, reflecting their potential applicability in clinical medicine.
Wenhan Dou, Shanshan Gao 0003, Deqian Mao, Honghao Dai, Chenhao Zhang 0001, Yuanfeng Zhou
Comput. Animat. Virtual Worlds6
2022 Grayscale self-adjusting network with weak feature enhancement for 3D lumbar anatomy segmentation
Jinhua Liu 0003, Zhiming Cui 0001, Christian Desrosiers, Shuyi Lu, Yuanfeng Zhou
Medical Image Anal.5
2022 DHNet: Salient Object Detection With Dynamic Scale-Aware Learning and Hard-Sample Refinement
abstract
During the annotation procedure of salient object detection, researchers usually locate the approximate location of the salient objects first and then process the pixels that need to be finely annotated. Following this idea, we find that the existing methods have limited exploration for solving the problem of positioning salient objects. Furthermore, no effective solution has been proposed for the hard-sample problem related to this task. Therefore, we propose dynamic scale-aware learning to learn dynamic scale weights that vary with different images to solve the first problem. Second, we design a dense sampling strategy for hard samples to construct a graph representation with samples from different classes and different confidence levels. Then, we achieve targeted feature aggregation based on the constructed graph with the help of the graph attention mechanism. We conduct extensive experiments on five benchmark datasets using comprehensive evaluation metrics. The results show that our method outperforms the current state-of-the-art approaches.
Chenhao Zhang 0001, Shanshan Gao 0003, Deqian Mao, Yuanfeng Zhou
IEEE Trans. Circuits Syst. Video Technol.4
2022 Fast Generation of Superpixels With Lattice Topology
abstract
Serving as an essential step for many applications of image processing, superpixel generation has attracted a lot of attentions. Most existing superpixel generation algorithms focus on the boundary adherence and compactness of the superpixels, but ignore the topological consistency between the superpixels, which severely limites their applications in the subsequent tasks, especially in the CNN based image processing tasks. In this paper, we present a fast lattice superpixel generation algorithm, which can generate superpixels with lattice topology like the original pixels. We also propose a local similarity loss function to improve the segmentation accuracy of the generated lattice superpixels. The whole algorithm is parallelly implemented on GPU. We perform extensive experiments on three datasets (i.e., BSDS500, NYUv2 and VOC) to verify the efficacy of our algorithm. The experimental results show that our method achieves competitive results compared to the state-of-the-art methods.
Yuanfeng Zhou, Yunfeng Zhang 0001, Caiming Zhang 0001
IEEE Trans. Image Process.2
2022 GDR-Net: A Geometric Detail Recovering Network for 3D Scanned Objects
abstract
This article addresses the problem of mesh super-resolution such that the geometry details which are not well represented in the low-resolution models can be recovered and well represented in the generated high-quality models. The main challenges of this problem are the nonregularity of 3D mesh representation and the high complexity of 3D shapes. We propose a deep neural network called GDR-Net to solve this ill-posed problem, which resolves the two challenges simultaneously. First, to overcome the nonregularity, we regress a displacement in radial basis function parameter space instead of the vertex-wise coordinates in the euclidean space. Second, to overcome the high complexity, we apply the detail recovery process to small surface patches extracted from the input surface and obtain the overall high-quality mesh by fusing the refined surface patches. To train the network, we constructed a dataset composed of both real-world and synthetic scanned models, including high/low-quality pairs. Our experimental results demonstrate that GDR-Net works well for general models and outperforms previous methods for recovering geometric details.
Wanquan Feng, Juyong Zhang, Yuanfeng Zhou, Shi-Qing Xin
IEEE Trans. Vis. Comput. Graph.3
2022 Probability driven approach for point cloud registration of indoor scene
Shanshan Gao 0003, Shi-Qing Xin, Yuanfeng Zhou
Vis. Comput.4
2021 Simplicity Driven Edge Refinement and Color Reconstruction in Image Vectorization
Junhao Zhao, Shi-Qing Xin, Shuang-Min Chen, Yuanfeng Zhou, Changhe Tu, Wenping Wang 0001
CGI5
2021 Context-Aware Virtual Adversarial Training for Anatomically-Plausible Segmentation
Ping Wang 0016, Jizong Peng, Marco Pedersoli, Yuanfeng Zhou, Caiming Zhang 0001, Christian Desrosiers
MICCAI (1)4
2021 Multi-Task Joint Learning of 3D Keypoint Saliency and Correspondence Estimation
Guangshun Wei, Long Ma 0009, Chen Wang 0054, Christian Desrosiers, Yuanfeng Zhou
Comput. Aided Des.5
2021 Compact joints encoding for skeleton-based dynamic hand gesture recognition
Yangke Li, Dongyang Ma, Yuhang Yu, Guangshun Wei, Yuanfeng Zhou
Comput. Graph.5
2021 Multi-scale joint feature network for micro-expression recognition
abstract
Micro-expression recognition is a substantive cross-study of psychology and computer science, and it has a wide range of applications (e.g., psychological and clinical diagnosis, emotional analysis, criminal investigation, etc.). However, the subtle and diverse changes in facial muscles make it difficult for existing methods to extract effective features, which limits the improvement of micro-expression recognition accuracy. Therefore, we propose a multi-scale joint feature network based on optical flow images for micro-expression recognition. First, we generate an optical flow image that reflects subtle facial motion information. The optical flow image is then fed into the multi-scale joint network for feature extraction and classification. The proposed joint feature module (JFM) integrates features from different layers, which is beneficial for the capture of micro-expression features with different amplitudes. To improve the recognition ability of the model, we also adopt a strategy for fusing the feature prediction results of the three JFMs with the backbone network. Our experimental results show that our method is superior to state-of-the-art methods on three benchmark datasets (SMIC, CASME II, and SAMM) and a combined dataset (3DB).
Guangshun Wei, Yuanfeng Zhou
Comput. Vis. Media4
2021 Image smoothing based on global sparsity decomposition and a variable parameter
abstract
Smoothing images, especially with rich texture, is an important problem in computer vision. Obtaining an ideal result is difficult due to complexity, irregularity, and anisotropicity of the texture. Besides, some properties are shared by the texture and the structure in an image. It is a hard compromise to retain structure and simultaneously remove texture. To create an ideal algorithm for image smoothing, we face three problems. For images with rich textures, the smoothing effect should be enhanced. We should overcome inconsistency of smoothing results in different parts of the image. It is necessary to create a method to evaluate the smoothing effect. We apply texture pre-removal based on global sparse decomposition with a variable smoothing parameter to solve the first two problems. A parametric surface constructed by an improved Bessel method is used to determine the smoothing parameter. Three evaluation measures: edge integrity rate, texture removal rate, and gradient value distribution are proposed to cope with the third problem. We use the alternating direction method of multipliers to complete the whole algorithm and obtain the results. Experiments show that our algorithm is better than existing algorithms both visually and quantitatively. We also demonstrate our method's ability in other applications such as clip-art compression artifact removal and content-aware image manipulation.
Xiang Ma 0006, Xuemei Li 0001, Yuanfeng Zhou, Caiming Zhang 0001
Comput. Vis. Media3
2021 TSegNet: An efficient and accurate tooth segmentation network on 3D dental model
Zhiming Cui 0001, Changjian Li 0001, Nenglun Chen, Guodong Wei, Runnan Chen, Yuanfeng Zhou, Dinggang Shen, Wenping Wang 0001
Medical Image Anal.6
2021 Self-paced and self-consistent co-training for semi-supervised image segmentation
Ping Wang 0016, Jizong Peng, Marco Pedersoli, Yuanfeng Zhou, Caiming Zhang 0001, Christian Desrosiers
Medical Image Anal.4
2021 Convex and Compact Superpixels by Edge- Constrained Centroidal Power Diagram
abstract
Superpixel segmentation, as a central image processing task, has many applications in computer vision and computer graphics. Boundary alignment and shape compactness are leading indicators to evaluate a superpixel segmentation algorithm. Furthermore, convexity can make superpixels reflect more geometric structures in images and provide a more concise over-segmentation result. In this paper, we consider generating convex and compact superpixels while satisfying the constraints of adhering to the boundary as far as possible. We formulate the new superpixel segmentation into an edge-constrained centroidal power diagram (ECCPD) optimization problem. In the implementation, we optimize the superpixel configurations by repeatedly performing two alternative operations, which include site location updating and weight updating through a weight function defined by image features. Compared with existing superpixel methods, our method can partition an image into fully convex and compact superpixels with better boundary adherence. Extensive experimental results show that our approach outperforms existing superpixel segmentation methods in boundary alignment and compactness for generating convex superpixels.
Dongyang Ma, Yuanfeng Zhou, Shi-Qing Xin, Wenping Wang 0001
IEEE Trans. Image Process.2
2021 Top-Down Shape Abstraction Based on Greedy Pole Selection
abstract
Motivated by the fact that the medial axis transform is able to encode the shape completely, we propose to use as few medial balls as possible to approximate the original enclosed volume by the boundary surface. We progressively select new medial balls, in a top-down style, to enlarge the region spanned by the existing medial balls. The key spirit of the selection strategy is to encourage large medial balls while imposing given geometric constraints. We further propose a speedup technique based on a provable observation that the intersection of medial balls implies the adjacency of power cells (in the sense of the power crust).We further elaborate the selection rules in combination with two closely related applications. One application is to develop an easy-to-use ball-stick modeling system that helps non-professional users to quickly build a shape with only balls and wires, but any penetration between two medial balls must be suppressed. The other application is to generate porous structures with convex, compact (with a high isoperimetric quotient) and shape-aware pores where two adjacent spherical pores may have penetration as long as the mechanical rigidity can be well preserved.
Zhiyang Dou, Shi-Qing Xin, Rui Xu 0016, Jian Xu 0023, Yuanfeng Zhou, Shuang-Min Chen, Wenping Wang 0001, Xiuyang Zhao, Changhe Tu
IEEE Trans. Vis. Comput. Graph.5
2020 Computing Smooth Quasi-geodesic Distance Field (QGDF) with Quadratic Programming
Luming Cao, Junhao Zhao, Jian Xu 0023, Shuang-Min Chen, Guozhu Liu, Shi-Qing Xin, Yuanfeng Zhou, Ying He 0001
Comput. Aided Des.7
2020 Skeletonization via dual of shape segmentation
Jingliang Cheng, Shuang-Min Chen, Guozhu Liu, Shi-Qing Xin, Lin Lu 0001, Yuanfeng Zhou, Changhe Tu
Comput. Aided Geom. Des.7
2020 Robustly computing restricted Voronoi diagrams (RVD) on thin-plate models
Shi-Qing Xin, Changhe Tu, Dong-Ming Yan 0001, Yuanfeng Zhou, Caiming Zhang 0001
Comput. Aided Geom. Des.5
2020 SRF-Net: Spatial Relationship Feature Network for Tooth Point Cloud Classification
abstract
Abstract 3D scanned point cloud data of teeth is popular used in digital orthodontics. The classification and semantic labelling for point cloud of each tooth is a key and challenging task for planning dental treatment. Utilizing the priori ordered position information of tooth arrangement, we propose an effective network for tooth model classification in this paper. The relative position and the adjacency similarity feature vectors are calculated for tooth 3D model, and combine the geometric feature into the fully connected layers of the classification training task. For the classification of dental anomalies, we present a dental anomalies processing method to improve the classification accuracy. We also use FocalLoss as the loss function to solve the sample imbalance of wisdom teeth. The extensive evaluations, ablation studies and comparisons demonstrate that the proposed network can classify tooth models accurately and automatically and outperforms state‐of‐the‐art point cloud classification methods.
Guangshun Wei, Yuanfeng Zhou, Shi-Qing Xin, Wenping Wang 0001
Comput. Graph. Forum3
2020 Coarse to Fine: Weak Feature Boosting Network for Salient Object Detection
abstract
Abstract Salient object detection is to identify objects or regions with maximum visual recognition in an image, which brings significant help and improvement to many computer visual processing tasks. Although lots of methods have occurred for salient object detection, the problem is still not perfectly solved especially when the background scene is complex or the salient object is small. In this paper, we propose a novel Weak Feature Boosting Network (WFBNet) for the salient object detection task. In the WFBNet, we extract the unpredictable regions (low confidence regions) of the image via a polynomial function and enhance the features of these regions through a well‐designed weak feature boosting module (WFBM). Starting from a coarse saliency map, we gradually refine it according to the boosted features to obtain the final saliency map, and our network does not need any post‐processing step. We conduct extensive experiments on five benchmark datasets using comprehensive evaluation metrics. The results show that our algorithm has considerable advantages over the existing state‐of‐the‐art methods.
Chenhao Zhang 0001, Shanshan Gao 0003, Yuanfeng Zhou
Comput. Graph. Forum5
2020 Skeletal saliency map computation based on projection symmetry analysis
Shi-Qing Xin, Shanshan Gao 0003, Yuanfeng Zhou
Graph. Model.4
2020 Cervical cancer detection in cervical smear images using deep pyramid inference with refinement and spatial-aware booster
abstract
With the development of artificial intelligence and image processing technology, more and more intelligent diagnosis technologies are used in cervical cancer screening. Among them, the detection of cervical lesions by thin liquid‐based cytology is the most common method for cervical cancer screening. At present, most cervical cancer detection algorithms use the object detection technology of natural images, and often only minor modifications are made while ignoring the specificity of the complex application scenario of cervical lesions detection in cervical smear images. In this study, the authors combine the domain knowledge of cervical cancer detection and the characteristics of pathological cells to design a network and propose a booster for cervical cancer detection (CCDB). The booster mainly consists of two components: the refinement module and the spatial‐aware module. The characteristics of cancer cells are fully considered in the booster, and the booster is light and transplantable. As far as the authors know, they are the first to design a CCDB according to the characteristics of cervical cancer cells. Compared with baseline (Retinanet), the sensitivity at four false positives per image and average precision of the proposed method are improved by 2.79 and 7.2%, respectively.
Dongyang Ma, Jinhua Liu 0003, Yuanfeng Zhou
IET Image Process.4
2020 Att-MoE: Attention-based Mixture of Experts for nuclear and cytoplasmic segmentation
Jinhua Liu 0003, Christian Desrosiers, Yuanfeng Zhou
Neurocomputing3
2020 Two-Stream Temporal Convolutional Networks for Skeleton-Based Human Action Recognition
Jin-Gong Jia, Yuanfeng Zhou, Xing-Wei Hao, Feng Li 0002, Christian Desrosiers, Caiming Zhang 0001
J. Comput. Sci. Technol.2
2019 Field-aligned Quadrangulation for Image Vectorization
abstract
Abstract Image vectorization is an important yet challenging problem, especially when the input image has rich content. In this paper, we develop a novel method for automatically vectorizing natural images with feature‐aligned quad‐dominant meshes. Inspired by the quadrangulation methods in 3D geometry processing, we propose a new directional field optimization technique by encoding the color gradients, sidestepping the explicit computing of salient image features. We further compute the anisotropic scales of the directional field by accommodating the distance among image features. Our method is fully automatic and efficient, which takes only a few seconds for a 400×400 image on a normal laptop. We demonstrate the effectiveness of the proposed method on various image editing applications.
Guangshun Wei, Yuanfeng Zhou, Xifeng Gao, Shi-Qing Xin, Ying He 0001
Comput. Graph. Forum2
2019 Texture Relative Superpixel Generation With Adaptive Parameters
abstract
Superpixel generation, which is an essential step in many image processing applications, has attracted increasing attention from researchers. In this paper, we present an efficient flooding-based superpixel generation algorithm that generates compact and highly boundary adherent superpixels. In particular, by considering various superpixel properties, we measure the similarities between image pixels by proposing a new distance metric that combines various image features (e.g., colors, spatial locations, neighbor information, and texture features). To control the relative significance of these image features, we acquire the weights of image features (e.g., colors, texture features, and neighbor information of pixels) through a neural network. Then, the final superpixels are obtained through a greedy optimization that considers both the current superpixel and its neighboring superpixels. We perform extensive experiments on two datasets to verify the efficacy of our algorithm. The results show that our algorithm has considerable advantages over existing state-of-the-art methods, particularly regarding the compactness of the resulting superpixels.
Yuanfeng Zhou, Zhonggui Chen, Caiming Zhang 0001
IEEE Trans. Multim.2
2019 Constructing 3D Self-Supporting Surfaces with Isotropic Stress Using 4D Minimal Hypersurfaces of Revolution
abstract
This article presents a new computational framework for constructing 3D self-supporting surfaces with isotropic stress. Inspired by the self-supporting property of catenary and the fact that catenoid (the surface of revolution of the catenary curve) is a minimal surface, we discover the relation between 3D self-supporting surfaces and 4D minimal hypersurfaces (which are 3-manifolds). Lifting the problem into 4D allows us to convert gravitational forces into tensions and reformulate the equilibrium problem to total potential energy minimization, which can be solved using a variational method. We prove that the hyper-generatrix of a 4D minimal hyper-surface of revolution is a 3D self-supporting surface, implying that constructing a 3D self-supporting surface is equivalent to volume minimization. We show that the energy functional is simply the surface’s gravitational potential energy, which in turn can be converted into a surface reconstruction problem with mean curvature constraint. Armed with our theoretical findings, we develop an iterative algorithm to construct 3D self-supporting surfaces from triangle meshes. Our method guarantees convergence and can produce near-regular triangle meshes, thanks to a local mesh refinement strategy similar to centroidal Voronoi tessellation. It also allows users to tune the geometry via specifying either the zero potential surface or its desired volume. We also develop a finite element method to verify the equilibrium condition on 3D triangle meshes. The existing thrust network analysis methods discretize both geometry and material by approximating the continuous stress field through uniaxial singular stresses, making them an ideal tool for analysis and design of beam structures. In contrast, our method works on piecewise linear surfaces with continuous material. Moreover, our method does not require the 3D-to-2D projection, therefore it also works for both height and non-height fields.
Long Ma 0009, Ying He 0001, Qian Sun 0003, Yuanfeng Zhou, Caiming Zhang 0001, Wenping Wang 0001
ACM Trans. Graph.4
2018 Lightweight preprocessing and fast query of geodesic distance via proximity graph
Shi-Qing Xin, Wenping Wang 0001, Ying He 0001, Yuanfeng Zhou, Shuang-Min Chen, Changhe Tu, Zhenyu Shu
Comput. Aided Des.4
2018 2D skeleton extraction based on heat equation
Fengyi Gao, Guangshun Wei, Shi-Qing Xin, Shanshan Gao 0003, Yuanfeng Zhou
Comput. Graph.5
2017 Superpixels by Bilateral Geodesic Distance
abstract
We present a novel superpixel generation algorithm based on a new definition of geodesic distance, called bilateral geodesic distance. In contrast to the traditional geodesic distance, the new bilateral geodesic distance of two pixels considers the distance between their positions as well as their color difference. Superpixel generation is essentially a problem of clustering image pixels with respect to a set of properly selected seeds. We first use an adaptive hexagonal subdivision method to determine the initial seed-based image gradient. Then, we use the bilateral geodesic distance to measure the similarity between the pixels and the seeds. We apply an improved fast marching method to generate superpixels’ contour regions with the expansion velocities dependent on a new gradient formulation that depends on the seeds’ properties. The experimental results indicate that our algorithm is not only much faster than the structure-based method, which uses conventional geodesic distance, but also outperforms the existing methods in terms of region compactness and region boundary regularity.
Yuanfeng Zhou, Wenping Wang 0001, Yilong Yin, Caiming Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2017 Superpixels of RGB-D Images for Indoor Scenes Based on Weighted Geodesic Driven Metric
abstract
Serving as a key step for applications of image processing, superpixel generation has been attracting increasing attention. RGB-D images are used pervasively in scenes reconstruction and representation, benefiting from their contained depth data. In this paper, we present a novel framework for generating superpixels focus on RGB-D images of indoor scenes, based on a weighted geodesic driven metric that combines both color and geometric information. In particular, taking into account the unique structures of indoor scenarios, we first denoise the given RGB-D image, and construct the corresponding triangular mesh. A new weighted geodesic driven metric is defined by introducing a weight function constrained with normal vectors and colors. Under this metric, an energy function is defined to measure our over-segmentation of the triangular mesh, by optimizing which, we can acquire an optimal over-segmentation of the triangular mesh with object boundaries respected, such that vertices in each sub-region have similar geometric structures and color intensities. Re-mapping the over-segmentation of the triangular mesh to the RGB-D image results in desired superpixels. We perform extensive experiments on a large-scale database of RGB-D images to verify the efficacy of our algorithm. The results show that our algorithm has considerable advantages over the existing state-of-the-art methods.
Yuanfeng Zhou, Feng Li 0002, Caiming Zhang 0001
IEEE Trans. Vis. Comput. Graph.2
2017 Two-dimensional shape retrieval using the distribution of extrema of Laplacian eigenfunctions
Dongmei Niu, Peer-Timo Bremer, Peter Lindstrom 0001, Bernd Hamann, Yuanfeng Zhou, Caiming Zhang 0001
Vis. Comput.5
2016 A framework for modeling high quality tension-determined surfaces
Long Ma 0009, Yuanfeng Zhou, Hao Pan 0001, Caiming Zhang 0001
Comput. Graph.2
2016 Non-local feature back-projection for image super-resolution
abstract
Image super‐resolution (SR) for a single low‐resolution image is an important and challenging task in image processing. In this study, the authors propose a novel non‐local feature back‐projection method for image SR, which can effectively reduce jaggy and ringing artefacts common, in general, iterative back‐projection (IBP) method. In their method, the objective high‐resolution (HR) image is obtained by projecting reconstructed errors back to HR image iteratively. To optimise the initial HR image and constrain anisotropic errors propagation during IBP process, an efficient non‐local feature interpolation algorithm is designed. Specially, edge information is used as constraints to make the interpolation surface preserve better shape. Furthermore, as post‐processing, non‐local similarities are utilised to remove noise and irregularities induced by errors propagation. Experimental results show that their method achieves better performance than state‐of‐the‐art methods in terms of both quantitative metrics and visual qualities.
Xin Zhang 0079, Xuemei Li 0001, Yuanfeng Zhou, Caiming Zhang 0001
IET Image Process.4
2016 Adaptive sparse coding on PCA dictionary for image denoising
Caiming Zhang 0001, Qiang Guo 0003, Yuanfeng Zhou
Vis. Comput.5
2015 A nonlocal gradient concentration method for image smoothing
abstract
It is challenging to consistently smooth natural images, yet smoothing results determine the quality of a broad range of applications in computer vision. To achieve consistent smoothing, we propose a novel optimization model making use of the redundancy of natural images, by defining a nonlocal concentration regularization term on the gradient. This nonlocal constraint is carefully combined with a gradient-sparsity constraint, allowing details throughout the whole image to be removed automatically in a data-driven manner. As variations in gradient between similar patches can be suppressed effectively, the new model has excellent edge preserving, detail removal, and visual consistency properties. Comparisons with state-of-the-art smoothing methods demonstrate the effectiveness of the new method. Several applications, including edge manipulation, image abstraction, detail magnification, and image resizing, show the applicability of the new method.
Caiming Zhang 0001, Qiang Guo 0003, Yuanfeng Zhou
Comput. Vis. Media4
2015 Efficient tetrahedral mesh generation based on sampling optimization
abstract
Abstract We present a heuristic approach to tetrahedral mesh generation for implicit closed surfaces. It consists of a surface sampling step and a volume sampling step that both work in a unified optimization framework. First, high‐quality isotropic samplings as well as a triangular mesh on the surface are generated. Then uniform volume samplings are determined by optimizing the point distribution inside the closed surface domain. Finally, the tetrahedral mesh is easily obtained by constrained Delaunay triangulation. Experimental results show that the new method can generate ideal tetrahedral meshes for closed implicit surfaces efficiently that are Delaunay based. Our method has the advantage of high efficiency and nice performance at surface boundaries. Copyright © 2015 John Wiley & Sons, Ltd.
Yuanfeng Zhou, Caiming Zhang 0001, Pengbo Bo
Comput. Animat. Virtual Worlds1
2014 Flooding based superpixels generation with color, compactness and smoothness constraints
abstract
Superpixel generation is widely used in image segmentation. In this paper, we present an efficient flooding based superpixel generation algorithm. A new distance metric is defined for estimating pixels and seeds' similarity with COLOR, COMPACTNESS and SMOOTHNESS constraints. A seeds update strategy based on Lloyd's algorithm is adopted for optimizing seeds and superpixels' contour regions. Experimental results show our algorithm outperforms the existing methods. Boundaries of superpixels generated by our algorithm can fit the original image boundaries better.
Yuanfeng Zhou, Caiming Zhang 0001
ICIP2
2014 Surface Interpolation to Image with Edge Preserving
abstract
Based on the assumption that low-resolution image is sampled from an original scene approximated by piecewise polynomial surface, this paper proposes two different constraints to make the fitting surface of the original scene preserve the image's edge characteristics. We first construct a cubic parametric curve to approximate the edge in each local area and obtain an auxiliary pixel set by re-sampling the curve, then introduce a weight function to accord pixels different degrees of effects on surface reconstruction. By re-sampling the constructed surface, the enlarged image can be easily obtained. Extensive experimental results on various types of low-resolution images demonstrate that our method produces generally better results both in terms of quantitative evaluation and subjective visual quality.
Caiming Zhang 0001, Yuanfeng Zhou, Xuemei Li 0001
ICPR3
2014 Mesh resizing based on hierarchical saliency detection
Shixiang Jia, Caiming Zhang 0001, Xuemei Li 0001, Yuanfeng Zhou
Graph. Model.4
2014 Robust multi-level partition of unity implicits from triangular meshes
abstract
ABSTRACT This paper presents a new robust multi‐level partition of unity (MPU) method, which constructs an implicit surface from a triangular mesh via the new error metric between the mesh and the implicit surface. The new error metric employs a weighted function of inner points and vertices of a triangle to fit an implicit surface, which can control the approximation error between the surface and vertices of the triangle. Furthermore, it is applied to the MPU method by utilizing the dual graph of a triangular mesh, and the general quadric implicit surface is used for surface representation. Compared with the MPU method, the new method generates fewer subdivision cells with the same approximation error and performs more steadily especially when given triangular mesh with fewer vertices. Copyright © 2013 John Wiley & Sons, Ltd.
Yuanfeng Zhou, Caiming Zhang 0001, Xuemei Li 0001
Comput. Animat. Virtual Worlds2
2013 Fitting Multiple Curves to Point Clouds with Complicated Topological Structures
abstract
We present an automatic method for fitting multiple B-spline curves to unorganized planar points. The method works on point clouds which have complicated topological structures and a single curve is insufficient for fitting the shape. A divide-and-merge algorithm is developed for dividing the unorganized data points into several groups while each group represents a smooth curve. Each point group is then fitted with a B-spline curve by the SDM method. Our algorithm also sets up automatically the control polygon of initial B-spline curves. Experiments demonstrate the capability of the presented algorithm in accurate reconstruction of topological structures of point clouds.
Dongfang Zhu, Pengbo Bo, Yuanfeng Zhou, Caiming Zhang 0001, Kuanquan Wang
CAD/Graphics3
2012 Global Contrast of Superpixels Based Salient Region Detection
Caiming Zhang 0001, Yuanfeng Zhou, Yu Wei 0004
CVM3
2010 Fast Updating of Delaunay Triangulation of Moving Points by Bi-cell Filtering
abstract
Abstract Updating a Delaunay triangulation when data points are slightly moved is the bottleneck of computation time in variational methods for mesh generation and remeshing. Utilizing the connectivity coherence between two consecutive Delaunay triangulations for computation speedup is the key to solving this problem. Our contribution is an effective filtering technique that confirms most bi‐cells whose Delaunay connectivities remain unchanged after the points are perturbed. Based on bi‐cell flipping, we present an efficient algorithm for updating two‐dimensional and three‐dimensional Delaunay triangulations of dynamic point sets. Experimental results show that our algorithm outperforms previous methods.
Yuanfeng Zhou, Feng Sun 0006, Wenping Wang 0001, Jiaye Wang, Caiming Zhang 0001
Comput. Graph. Forum1
2007 A Quasi-Laplacian Smoothing Approach on Arbitrary Triangular Meshes
abstract
A new method for smoothing triangular meshes is presented. Mean curvature normal is used to define a Quasi-Laplacian for smoothing inner vertices at a local region. Vertices are moved along the normal direction in a more appropriate velocity which can make mesh smoothing and shape preserving harmonizing well. For the boundary vertices, a new method for estimating the mean curvature normal is presented, so that for an arbitrary triangular mesh, the inner and the boundary vertices can be smoothed by the same smoothing process. Features of the original mesh can be preserved by the weighted mean curvature normal restriction of the neighbors of one vertex effectively. Experiments of comparison between the new method and previous methods are included in this paper.
Yuanfeng Zhou, Caiming Zhang 0001, Shanshan Gao 0003
CAD/Graphics1