Ruofeng Tong 0001

dblp:11/914-1 · also Ruo-feng Tong 0001 · DBLP profile ↗
← Back
127ranked-venue papers
5as first author
53since 2021 · last 2026
0000-0002-8167-5354ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 90 · 4 first-author · 39 since 2021Applied, interdisciplinary, general and emerging computing · 28 · 1 first-author · 15 since 2021Artificial intelligence and machine learning · 14 · 12 since 2021Human-computer interaction and ubiquitous computing · 14 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 CADDesigner: Conceptual CAD model generation with a general-purpose agent
Fengxiao Fan, Jingzhe Ni, Xiaolong Yin, Qiang Zou 0007, Ruofeng Tong 0001, Min Tang 0001
Comput. Aided Des.7
2026 MidSurfer: Efficient Mid-Surface Abstraction from Variable Thin-Walled Models
Xinhang Zhou, Ruofeng Tong 0001, Min Tang 0001
Comput. Aided Des.4
2026 gMidSurf: Hierarchical GPU-based mid-surface abstraction for thin-walled CAD models
Xinhang Zhou, Ruofeng Tong 0001, Min Tang 0001
Comput. Aided Des.5
2026 RLCAD: Reinforcement learning training gym for revolution involved CAD command sequence generation
Xiaolong Yin, Jiahang Shen, Jingzhe Ni, Ruofeng Tong 0001, Min Tang 0001
Comput. Aided Des.6
2026 Multimodal Graph Learning With Multi-Hypergraph Reasoning Networks for Focal Liver Lesion Classification in Multimodal Magnetic Resonance Imaging
abstract
Multimodal magnetic resonance imaging (MRI) is instrumental in differentiating liver lesions. The major challenge involves modeling reliable connections and simultaneously learning complementary information across various MRI sequences. While previous studies have primarily focused on multimodal integration in a pair-wise manner using few modalities, our research seeks to advance a more comprehensive understanding of interaction modeling by establishing complex high-order correlations among the diverse modalities in multimodal MRI. In this paper, we introduce a multimodal graph learning with multi-hypergraph reasoning network to capture the full spectrum of both pair-wise and group-wise relationships among different modalities. Specifically, a weight-shared encoder extracts features from regions of interest (ROI) images across all modalities. Subsequently, a collection of uniform hypergraphs are constructed with varying vertex configurations, allowing for the modeling of not only pair-wise correlations but also the high-order collaborations for relational reasoning. Following information propagation through the hypergraph message passing, adaptive intra-modality fusion module is proposed to effectively fuse feature representations from different hypergraphs of the same modality. Finally, all refined features are concatenated to prepare for the classification task. Our experimental evaluations, including focal liver lesions classification using the LLD-MMRI2023 dataset and early recurrence prediction of hepatocellular carcinoma using our internal datasets, demonstrate that our method significantly surpasses the performance of existing approaches, indicating the effectiveness of our model in handling both pair-wise and group-wise interactions across multiple modalities.
Shaocong Mo, Lanfen Lin, Ruofeng Tong 0001, Fang Wang 0030, Qingqing Chen 0001, Wenbin Ji, Yinhao Li 0002, Hongjie Hu, Yen-Wei Chen 0001
IEEE J. Biomed. Health Informatics4
2026 GPU-accelerated Certified Hausdorff Distance Between Triangle Meshes
abstract
Computing the directed Hausdorff distance between two triangle meshes is a fundamental operation in geometry processing and simulation. While existing certified branch-and-bound (B&B) methods are efficient for well-separated geometry, they can become prohibitively expensive on large models under tight tolerances and near-zero distance configurations where pruning is limited. We present a GPU-accelerated certified B&B algorithm that explicitly maintains enclosing lower and upper bounds on the directed Hausdorff distance and terminates once their normalized gap, measured with respect to the bounding-box diagonal of the source mesh, meets a user-prescribed tolerance. To map the inherently prioritized search to SIMT (single-instruction, multiple-thread) hardware, we replace priority queues and recursion with a sorted, double-buffered wavefront pipeline built from bulk-parallel worklists for bound evaluation, culling, subdivision, and compaction. To mitigate loose bounds on thin primitives while preserving predictable stream behavior, we introduce a fixed-cardinality adaptive subdivision scheme that selectively applies double longest-edge bisection. To remain robust in deep-refinement regimes, we add a resource-aware deferral mechanism that enforces a device-capacity invariant by prioritizing candidates likely to be culled while postponing expensive ones. Finally, we improve numerical robustness under FP32 (single precision) via triangle-local coordinate transforms and other conservative numerical safeguards, and enhance coherence by spatially ordering the active set and traversing the BVH (bounding volume hierarchy) in triangle packets. Under the same stopping tolerance, experiments on an NVIDIA RTX 5090 show that our GPU solver remains numerically consistent with the FP64 CPU baseline, with normalized cross-platform deviation below 0.01% in over 99.9% of cases. Our method achieves millisecond-scale runtimes capable of supporting interactive frame rates, even on models with millions of triangles. Across the comparison set, it delivers throughput speedups of 836× on the Thingi10K/TetWild benchmark ( A → B ) and 709× on the Thingi10K/Decimation benchmark. Code and data for this paper are available at https://github.com/fhp-transient/gpu-hausdorff.
Haopeng Fan, Min Tang 0001, Leonardo Sacht, Qiang Zou 0007, Ruofeng Tong 0001
ACM Trans. Graph.5
2025 Image-Based Virtual Try-On: A Survey
Dan Song 0006, Xuanpu Zhang, Weizhi Nie, Ruofeng Tong 0001, Mohan Kankanhalli, Anan Liu
Int. J. Comput. Vis.5
2025 3D Point Cloud Matching Based Selfie Generation for Chang'e-5
Xiao-Rui Chen, Meng-Fei Yang, Gao Zhang, Xiang-Jin Deng, Liu-Zhi Yang, Yun Yang 0001, Shou-Qian Sun, Ruofeng Tong 0001, Min Tang 0001
J. Comput. Sci. Technol.12
2025 Source-Free Model Adaptation for Unsupervised 3D Object Retrieval
abstract
With the explosive growth of 3D objects yet expensive annotation costs, unsupervised 3D object retrieval has become a popular but challenging research area. Existing labeled resources have been utilized to aid this task via transfer learning, which aligns the distribution of unlabeled data with the source one. However, the labeled resource are not always accessible due to the privacy disputes, limited computational capacity and other thorny restrictions. Therefore, we propose source-free model adaptation task for unsupervised 3D object management, which utilizes a pre-trained model to boost the performance with no access to source data and labels. Specifically, we compute representative prototypes to assume the source feature distribution, and design a bidirectional cumulative confidence-based adaptation strategy to adaptively align unlabeled samples towards prototypes. Subsequently, a dual-model distillation mechanism is proposed to generate source hypothesis for remedying the absence of ground-truth labels. The experiments on a cross-domain retrieval benchmark NTU-PSB (PSB-NTU) and a cross-modality retrieval benchmark MI3DOR also demonstrate the superiority of the proposed method even without access to raw data.
Dan Song 0006, Yiyao Wu, Yuting Ling, Diqiong Jiang, Ruofeng Tong 0001
IEEE Trans. Vis. Comput. Graph.6
2024 Combinatorial CNN-Transformer Learning with Manifold Constraints for Semi-supervised Medical Image Segmentation
abstract
Semi-supervised learning (SSL), as one of the dominant methods, aims at leveraging the unlabeled data to deal with the annotation dilemma of supervised learning, which has attracted much attentions in the medical image segmentation. Most of the existing approaches leverage a unitary network by convolutional neural networks (CNNs) with compulsory consistency of the predictions through small perturbations applied to inputs or models. The penalties of such a learning paradigm are that (1) CNN-based models place severe limitations on global learning; (2) rich and diverse class-level distributions are inhibited. In this paper, we present a novel CNN-Transformer learning framework in the manifold space for semi-supervised medical image segmentation. First, at intra-student level, we propose a novel class-wise consistency loss to facilitate the learning of both discriminative and compact target feature representations. Then, at inter-student level, we align the CNN and Transformer features using a prototype-based optimal transport method. Extensive experiments show that our method outperforms previous state-of-the-art methods on three public medical image segmentation benchmarks.
Huimin Huang 0002, Yawen Huang, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Yuexiang Li, Yefeng Zheng 0001
AAAI5
2024 Going Beyond Multi-Task Dense Prediction with Synergy Embedding Models
abstract
Multi-task visual scene understanding aims to leverage the relationships among a set of correlated tasks, which are solved simultaneously by embedding them within a unified network. However, most existing methods give rise to two primary concerns from a task-level perspective: (1) the lack of task-independent correspondences for distinct tasks, and (2) the neglect of explicit task-consensual dependencies among various tasks. To address these issues, we propose a novel synergy embedding models (SEM), which goes beyond multi-task dense prediction by leveraging two innovative designs: the intra-task hierarchy-adaptive module and the inter-task EM-interactive module. Specifically, the constructed intra-task module incorporates hierarchy-adaptive keys from multiple stages, enabling the efficient learning of specialized visual patterns with an optimal trade-off. In addition, the developed inter-task module learns interactions from a compact set of mutual bases among various tasks, benefiting from the expectation maximization (EM) algorithm. Extensive empirical evidence from two public benchmarks, NYUD-v2 and PASCAL-Context, demonstrates that SEM consistently outperforms state-of-the-art approaches across a range of metrics.
Huimin Huang 0002, Yawen Huang, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hao Zheng 0008, Yuexiang Li, Yefeng Zheng 0001
CVPR4
2024 IRLSG: Invariant Representation Learning for Single-Domain Generalization in Medical Image Segmentation
abstract
Single-domain generalization (SDG) can efficiently enhance model generalization while avoiding high annotation costs and privacy concerns. However, existing SDG methods are mainly based on data manipulation and meta-learning, which are not efficient enough due to the limited generalization performance and complex inference. In response to these challenges, we present a novel single domaininvariant representation learning approach for medical image segmentation, called IRLSG, with two appealing designs: (1) A Classscale Photo-metric Augmentation is first proposed to simulate unseen target domain that is sufficient in diversity and informativeness. After that, a Dual-Consistency Framework is further designed to constrain the consistency of intermediate features and segmentation results between the original and the augmented images, which helps to explore the domain-invariant representation. (2) A simple and effective Style Feature Whitening is designed to decouple and remove the domain-specific style from higher-order covariance statistics, which can further improve the modeling and generalization capability of the network. Experimental results on different benchmarks demonstrate that our IRLSG outperforms the current state-of-the-art methods in tackling single-domain generalization.
Ziwei Niu, Hao Sun 0013, Shuyi Ouyang, Shiao Xie, Yen-Wei Chen 0001, Ruofeng Tong 0001, Lanfen Lin
ICASSP6
2024 Semi-Decoupled 6D Pose Estimation via Multi-Modal Feature Fusion
abstract
The existing methods for 6D pose estimation based on RGB-D employ RGB images and observed point cloud derived from depth maps as input, then concurrently predicting both rotation and translation. However, rotation and translation possess distinct characteristics and scale ranges, and their simultaneous prediction can lead to mutual influence in the network parameter space. Additionally, the observed point cloud are susceptible to systematic noise and partial data loss, presenting challenges for the network to capture comprehensive object features. To address these issues, we propose the Semi-Decoupled 6D pose estimation via multi-modal feature fusion (SD6D). SD6D comprises a Multi-Modal Fusion Module and a Semi-Decoupled Prediction Module. The former dynamically fuses different modal data (RGB, depth, CAD model) based on their inter-modality correlations, aiding in establishing 2D-3D correspondences and addressing issues stemming from systematic noise and partial data loss. The latter semi-decouples the prediction of rotation and translation, predicting them separately based on their distinct characteristics. We conducted experiments on two popular benchmark datasets, which prove the superiority of our method.
Zhenhu Zhang, Xin Cao 0010, Xueying Qin, Ruofeng Tong 0001
ICASSP5
2024 FedDGP: Disentangling Global and Personal Models for Federated Learning
abstract
Federated learning (FL) aims to construct a global model by collaboratively training local models on data with diverse distributions, emphasizing the exchange of model parameters rather than raw data sharing. In the medical domain, achieving high performance in local models, strong generalization capabilities in the global model, and minimizing communication costs are all crucial. However, current federated learning methods struggle to concurrently optimize these three aspects. This paper introduces FedDGP, an innovative FL framework addressing these challenges. FedDGP comprises three key modules: Domain-guided Model Disentanglement (MD), Heterogeneous Aggregation (HA), and Reciprocal Iterative Training (RIT). MD divides a client model into two parts with different task objectives, enhancing both client and server model performance while avoiding issues like catastrophic forgetting. HA assigns reduced weight to underperforming models, limiting their impact. RIT minimizes communication costs by uploading only one client model’s parameters per round, facilitating local optimal model transfers. Extensive experiments on six domain classification datasets demonstrate FedDGP’s effectiveness, showcasing improved performance and reduced communication costs compared to existing approaches.
Zhenhu Zhang, Dan Song 0006, Jiahua Dong 0001, Ruofeng Tong 0001
ICME5
2024 gDist: Efficient Distance Computation between 3D Meshes on GPU
Wei Wang 0419, Ruofeng Tong 0001, Min Tang 0001
SIGGRAPH Asia3
2024 CTSN: Predicting cloth deformation for skeleton-based characters with a two-stream skinning network
abstract
We present a novel learning method using a two-stream network to predict cloth deformation for skeleton-based characters. The characters processed in our approach are not limited to humans, and can be other targets with skeleton-based representations such as fish or pets. We use a novel network architecture which consists of skeleton-based and mesh-based residual networks to learn the coarse features and wrinkle features forming the overall residual from the template cloth mesh. Our network may be used to predict the deformation for loose or tight-fitting clothing. The memory footprint of our network is low, thereby resulting in reduced computational requirements. In practice, a prediction for a single cloth mesh for a skeleton-based character takes about 7 ms on an nVidia GeForce RTX 3090 GPU. Compared to prior methods, our network can generate finer deformation results with details and wrinkles.
Yudi Li, Min Tang 0001, Yun Yang 0001, Ruofeng Tong 0001, Shuangcai Yang, Bailin An, Qilong Kou
Comput. Vis. Media4
2024 Rethinking Multiple Instance Learning for Whole Slide Image Classification: A Bag-Level Classifier is a Good Instance-Level Teacher
abstract
Multiple Instance Learning (MIL) has demonstrated promise in Whole Slide Image (WSI) classification. However, a major challenge persists due to the high computational cost associated with processing these gigapixel images. Existing methods generally adopt a two-stage approach, comprising a non-learnable feature embedding stage and a classifier training stage. Though it can greatly reduce memory consumption by using a fixed feature embedder pre-trained on other domains, such a scheme also results in a disparity between the two stages, leading to suboptimal classification accuracy. To address this issue, we propose that a bag-level classifier can be a good instance-level teacher. Based on this idea, we design Iteratively Coupled Multiple Instance Learning (ICMIL) to couple the embedder and the bag classifier at a low cost. ICMIL initially fixes the patch embedder to train the bag classifier, followed by fixing the bag classifier to fine-tune the patch embedder. The refined embedder can then generate better representations in return, leading to a more accurate classifier for the next iteration. To realize more flexible and more effective embedder fine-tuning, we also introduce a teacher-student framework to efficiently distill the category knowledge in the bag classifier to help the instance-level embedder fine-tuning. Intensive experiments were conducted on four distinct datasets to validate the effectiveness of ICMIL. The experimental results consistently demonstrated that our method significantly improves the performance of existing MIL backbones, achieving state-of-the-art results. The code and the organized datasets can be accessed by: https://github.com/Dootmaan/ICMIL/tree/confidence-based.
Hongyi Wang 0002, Luyang Luo, Fang Wang 0030, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hongjie Hu, Lanfen Lin, Hao Chen 0011
IEEE Trans. Medical Imaging4
2024 Knowledge Distillation-Based Domain-Invariant Representation Learning for Domain Generalization
abstract
Domain generalization (DG) aims to generalize the knowledge learned from multiple source domains to unseen target domains. Existing DG techniques can be subsumed under two broad categories, i.e., domain-invariant representation learning and domain manipulation. Nevertheless, it is extremely difficult to explicitly augment or generate the unseen target data. And when source domain variety increases, developing a domain-invariant model by simply aligning more domain-specific information becomes more challenging. In this paper, we propose a simple yet effective method for domain generalization, named Knowledge Distillation based Domain-invariant Representation Learning (KDDRL), that learns domain-invariant representation while encouraging the model to maintain domain-specific features, which recently turned out to be effective for domain generalization. To this end, our method incorporates multiple auxiliary student models and one student leader model to perform a two-stage distillation. In the first-stage distillation, each domain-specific auxiliary student treats the ensemble of other auxiliary students' predictions as a target, which helps to excavate the domain-invariant representation. Also, we present an error removal module to prevent the transfer of faulty information by eliminating incorrect predictions compared to the true labels. In the second-stage distillation, the student leader model with domain-specific features combines the domain-invariant representation learned from the group of auxiliary students to make the final prediction. Extensive experiments and in-depth analysis on popular DG benchmark datasets demonstrate that our KDDRL significantly outperforms the current state-of-the-art methods.
Ziwei Niu, Junkun Yuan, Jing Liu 0041, Yen-Wei Chen 0001, Ruofeng Tong 0001, Lanfen Lin
IEEE Trans. Multim.7
2023 ClassFormer: Exploring Class-Aware Dependency with Transformer for Medical Image Segmentation
abstract
Vision Transformers have recently shown impressive performances on medical image segmentation. Despite their strong capability of modeling long-range dependencies, the current methods still give rise to two main concerns in a class-level perspective: (1) intra-class problem: the existing methods lacked in extracting class-specific correspondences of different pixels, which may lead to poor object coverage and/or boundary prediction; (2) inter-class problem: the existing methods failed to model explicit category-dependencies among various objects, which may result in inaccurate localization. In light of these two issues, we propose a novel transformer, called ClassFormer, powered by two appealing transformers, i.e., intra-class dynamic transformer and inter-class interactive transformer, to address the challenge of fully exploration on compactness and discrepancy. Technically, the intra-class dynamic transformer is first designed to decouple representations of different categories with an adaptive selection mechanism for compact learning, which optimally highlights the informative features to reflect the salient keys/values from multiple scales. We further introduce the inter-class interactive transformer to capture the category dependency among different objects, and model class tokens as the representative class centers to guide a global semantic reasoning. As a consequence, the feature consistency is ensured with the expense of intra-class penalization, while inter-class constraint strengthens the feature discriminability between different categories. Extensive empirical evidence shows that ClassFormer can be easily plugged into any architecture, and yields improvements over the state-of-the-art methods in three public benchmarks.
Huimin Huang 0002, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hong Wang 0021, Yuexiang Li, Yawen Huang, Yefeng Zheng 0001
AAAI4
2023 SemiCVT: Semi-Supervised Convolutional Vision Transformer for Semantic Segmentation
abstract
Semi-supervised learning improves data efficiency of deep models by leveraging unlabeled samples to alleviate the reliance on a large set of labeled samples. These successes concentrate on the pixel-wise consistency by using convolutional neural networks (CNNs) but fail to address both global learning capability and class-level features for unlabeled data. Recent works raise a new trend that Transformer achieves superior performance on the entire feature map in various tasks. In this paper, we unify the current dominant Mean-Teacher approaches by reconciling intra-model and inter-model properties for semi-supervised segmentation to produce a novel algorithm, SemiCVT, that absorbs the quintessence of CNNs and Transformer in a comprehensive way. Specifically, we first design a parallel CNN-Transformer architecture (CVT) with introducing an intra-model local-global interaction schema (LGI) in Fourier domain for full integration. The inter-model class-wise consistency is further presented to complement the class-level statistics of CNNs and Transformer in a cross-teaching manner. Extensive empirical evidence shows that SemiCVT yields consistent improvements over the state-of-the-art methods in two public benchmarks.
Huimin Huang 0002, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Yuexiang Li, Hong Wang 0021, Yawen Huang, Yefeng Zheng 0001
CVPR4
2023 StyleIPSB: Identity-Preserving Semantic Basis of StyleGAN for High Fidelity Face Swapping
abstract
Recent researches reveal that StyleGAN can generate highly realistic images, inspiring researchers to use pretrained StyleGAN to generate high-fidelity swapped faces. However, existing methods fail to meet the expectations in two essential aspects of high-fidelity face swapping. Their results are blurry without pore-level details and fail to preserve identity for challenging cases. To overcome the above artifacts, we innovatively construct a series of identity-preserving semantic bases of StyleGAN (called StyleIPSB) in respect of pose, expression, and illumination. Each basis of StyleIPSB controls one specific semantic attribute and disentangles with the others. The StyleIPSB constrains style code in the subspace of W+ space to preserve pore-level details and gives us a novel tool for high-fidelity face swapping, and we propose a three-stage framework for face swapping with StyleIPSB. Firstly, we transform the target facial images' attributes to the source image. We learn the mapping from 3D Morphable Model (3DMM) parameters, which capture the prominent semantic variance, to the coordinates of StyleIPSB that show higher identity-preserving and fidelity. Secondly, to transform detailed attributes which 3DMM does not capture, we learn the residual attribute between the reenacted face and the target face. Finally, the face is blended into the background of the target image. Extensive results and comparisons demonstrate that StyleIPSB can effectively preserve identity and pore-level details. The results of face swapping can achieve state-of-the-art performance. We will release our code at https://github.com/a686432/StyleIPSB
Diqiong Jiang, Dan Song 0006, Ruofeng Tong 0001, Min Tang 0001
CVPR3
2023 SLViT: Scale-Wise Language-Guided Vision Transformer for Referring Image Segmentation
abstract
Referring image segmentation aims to segment an object out of an image via a specific language expression. The main concept is establishing global visual-linguistic relationships to locate the object and identify boundaries using details of the image. Recently, various Transformer-based techniques have been proposed to efficiently leverage long-range cross-modal dependencies, enhancing performance for referring segmentation. However, existing methods consider visual feature extraction and cross-modal fusion separately, resulting in insufficient visual-linguistic alignment in semantic space. In addition, they employ sequential structures and hence lack multi-scale information interaction. To address these limitations, we propose a Scale-Wise Language-Guided Vision Transformer (SLViT) with two appealing designs: (1) Language-Guided Multi-Scale Fusion Attention, a novel attention mechanism module for extracting rich local visual information and modeling global visual-linguistic relationships in an integrated manner. (2) An Uncertain Region Cross-Scale Enhancement module that can identify regions of high uncertainty using linguistic features and refine them via aggregated multi-scale features. We have evaluated our method on three benchmark datasets. The experimental results demonstrate that SLViT surpasses state-of-the-art methods with lower computational cost. The code is publicly available at: https://github.com/NaturalKnight/SLViT.
Shuyi Ouyang, Hongyi Wang 0002, Shiao Xie, Ziwei Niu, Ruofeng Tong 0001, Yen-Wei Chen 0001, Lanfen Lin
IJCAI5
2023 Iteratively Coupled Multiple Instance Learning from Instance to Bag Classifier for Whole Slide Image Classification
Hongyi Wang 0002, Luyang Luo, Fang Wang 0030, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hongjie Hu, Lanfen Lin, Hao Chen 0011
MICCAI (6)4
2023 Semi-Supervised Convolutional Vision Transformer with Bi-Level Uncertainty Estimation for Medical Image Segmentation
abstract
Semi-supervised learning (SSL) has attracted much attention in the field of medical image segmentation, which enables to alleviate the heavy burden of labelling pixel-wise annotation by extracting knowledge from unlabeled data. The existing methods basically benefit from the success of convolutional neural networks (CNNs) by keeping consistency of the predictions under small perturbations imposed on the networks or inputs. Two main concerns arise when learning such a paradigm: (1) CNNs tend to retain discriminative local features, neglecting global dependency and thus leading to inaccurate localization; (2) CNNs omit reliable feature-level and pixel-level information, resulting in sketchy pseudo-labels, especially around the confusing boundary. In this paper, we revisit the model of semi-supervised learning and develop a novel CNN-Transformer learning framework that allows for effective segmentation of medical images by producing complementary and reliable features and pseudo-label with bi-level uncertainty. Motivated by the uncertainty estimation to gain insight on feature discrimination, we explore the statistical and geometrical properties of features on network optimization and thus launching an alignment method in a more accurate and stable way. We attach equal significance to pixel-level uncertainty estimation for alleviating the influence of unreliable pseudo-labels in the training progress and advocating the reliability of predictions. Experimental results show that our method significantly surpasses existing semi-supervised approaches on two public medical image segmentation datasets.
Huimin Huang 0002, Yawen Huang, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Yuexiang Li, Yefeng Zheng 0001
ACM Multimedia5
2023 HSVLT: Hierarchical Scale-Aware Vision-Language Transformer for Multi-Label Image Classification
abstract
The task of multi-label image classification involves recognizing multiple objects within a single image. Considering both valuable semantic information contained in the labels and essential visual features presented in the image, tight visual-linguistic interactions play a vital role in improving classification performance. Moreover, given the potential variance in object size and appearance within a single image, attention to features of different scales can help to discover possible objects in the image. Recently, Transformer-based methods have achieved great success in multi-label image classification by leveraging the advantage of modeling long-range dependencies, but they have several limitations. Firstly, existing methods treat visual feature extraction and cross-modal fusion as separate steps, resulting in insufficient visual-linguistic alignment in the joint semantic space. Additionally, they only extract visual features and perform cross-modal fusion at a single scale, neglecting objects with different characteristics. To address these issues, we propose a Hierarchical Scale-Aware Vision-Language Transformer (HSVLT) with two appealing designs: (1)A hierarchical multi-scale architecture that involves a Cross-Scale Aggregation module, which leverages joint multi-modal features extracted from multiple scales to recognize objects of varying sizes and appearances in images. (2)Interactive Visual-Linguistic Attention, a novel attention mechanism module that tightly integrates cross-modal interaction, enabling the joint updating of visual, linguistic and multi-modal features. We have evaluated our method on three benchmark datasets. The experimental results demonstrate that HSVLT surpasses state-of-the-art methods with lower computational cost.
Shuyi Ouyang, Hongyi Wang 0002, Ziwei Niu, Zhenjia Bai, Shiao Xie, Ruofeng Tong 0001, Yen-Wei Chen 0001, Lanfen Lin
ACM Multimedia7
2023 D-Cloth: Skinning-based Cloth Dynamic Prediction with a Three-stage Network
abstract
Abstract We propose a three‐stage network that utilizes a skinning‐based model to accurately predict dynamic cloth deformation. Our approach decomposes cloth deformation into three distinct components: static, coarse dynamic, and wrinkle dynamic components. To capture these components, we train our three‐stage network accordingly. In the first stage, the static component is predicted by constructing a static skinning model that incorporates learned joint increments and skinning weight increments. Then, in the second stage, the coarse dynamic component is added to the static skinning model by incorporating serialized skeleton information. Finally, in the third stage, the mesh sequence stage refines the prediction by incorporating the wrinkle dynamic component using serialized mesh information. We have implemented our network and used it in a Unity game scene, enabling real‐time prediction of cloth dynamics. Our implementation achieves impressive prediction speeds of approximately 3.65ms using an NVIDIA GeForce RTX 3090 GPU and 9.66ms on an Intel i7‐7700 CPU. Compared to SOTA methods, our network excels in accurately capturing fine dynamic cloth deformations.
Yudi Li, Min Tang 0001, X. R. Chen, Yun Yang 0001, Ruofeng Tong 0001, Bailin An, Shuangcai Yang, Qilong Kou
Comput. Graph. Forum5
2023 Sphere Face Model: A 3D morphable model with hypersphere manifold latent space using joint 2D/3D training
abstract
3D morphable models (3DMMs) are generative models for face shape and appearance. Recent works impose face recognition constraints on 3DMM shape parameters so that the face shapes of the same person remain consistent. However, the shape parameters of traditional 3DMMs satisfy the multivariate Gaussian distribution. In contrast, the identity embeddings meet the hypersphere distribution, and this conflict makes it challenging for face reconstruction models to preserve the faithfulness and the shape consistency simultaneously. In other words, recognition loss and reconstruction loss can not decrease jointly due to their conflict distribution. To address this issue, we propose the Sphere Face Model (SFM), a novel 3DMM for monocular face reconstruction, preserving both shape fidelity and identity consistency. The core of our SFM is the basis matrix which can be used to reconstruct 3D face shapes, and the basic matrix is learned by adopting a two-stage training approach where 3D and 2D training data are used in the first and second stages, respectively. We design a novel loss to resolve the distribution mismatch, enforcing that the shape parameters have the hyperspherical distribution. Our model accepts 2D and 3D data for constructing the sphere face models. Extensive experiments show that SFM has high representation ability and clustering performance in its shape parameter space. Moreover, it produces high-fidelity face shapes consistently in challenging conditions in monocular face reconstruction. The code will be released at https://github.com/a686432/SIR
Diqiong Jiang, Yiwei Jin, Zhe Zhu, Yun Zhang 0024, Ruofeng Tong 0001, Min Tang 0001
Comput. Vis. Media6
2023 Struct2Hair: A hair shape descriptor for hairstyle modeling
abstract
Abstract In recent years, it becomes possible to extract hair information for hair reconstruction from multiple cameras or monocular camera. Using a single image as the input avoids the high cost setups and complex calibration compared to multiviewed reconstruction. Taking advantage of an extendible hairstyle database, this paper introduced Struct2Hair, a novel single‐viewed hair modelling approach by extracting hair shape descriptor (HSD). The HSD is defined as the fundamental structure‐aware feature, which is a combination of critical shapes in a hairstyle. A complete dataset of critical hair shapes is constructed from a known database of three‐dimensional (3D) hair models. We first analyze the input two‐dimensional (2D) image to extract the orientation information and 2D hair sketch automatically. The extracted information is then used to retrieve the corresponding critical shapes with optimization to build the robust HSD. Finally, the HSD constructs a weighted 3D hair orientation field to guide full‐head hair model generation. Our method can preserve local geometric features of hair and retain the whole shape of the hairstyle globally owing to the HSD, which will benefit further hair editing and stylization.
Wenshu Zhang, Yinyu Nie, Shihui Guo, Jian Chang 0001, Jian J. Zhang 0001, Ruofeng Tong 0001
Comput. Animat. Virtual Worlds6
2023 Adaptive Decomposition and Shared Weight Volumetric Transformer Blocks for Efficient Patch-Free 3D Medical Image Segmentation
abstract
High resolution (HR) 3D medical image segmentation is vital for an accurate diagnosis. However, in the field of medical imaging, it is still a challenging task to achieve a high segmentation performance with cost-effective and feasible computation resources. Previous methods commonly use patch-sampling to reduce the input size, but this inevitably harms the global context and decreases the model's performance. In recent years, a few patch-free strategies have been presented to deal with this issue, but either they have limited performance due to their over-simplified model structures or they follow a complicated training process. In this study, to effectively address these issues, we present Adaptive Decomposition (A-Decomp) and Shared Weight Volumetric Transformer Blocks (SW-VTB). A-Decomp can adaptively decompose features and reduce their spatial size, which greatly lowers GPU memory consumption. SW-VTB is able to capture long-range dependencies at a low cost with its lightweight design and cross-scale weight-sharing mechanism. Our proposed cross-scale weight-sharing approach enhances the network's ability to capture scale-invariant core semantic information in addition to reducing parameter numbers. By combining these two designs together, we present a novel patch-free segmentation framework named VolumeFormer. Experimental results on two datasets show that VolumeFormer outperforms existing patch-based and patch-free methods with a comparatively fast inference speed and relatively compact design.
Hongyi Wang 0002, Qingqing Chen 0001, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hongjie Hu, Lanfen Lin
IEEE J. Biomed. Health Informatics4
2023 Multi-Modal Tumor Segmentation With Deformable Aggregation and Uncertain Region Inpainting
abstract
Multi-modal tumor segmentation exploits complementary information from different modalities to help recognize tumor regions. Known multi-modal segmentation methods mainly have deficiencies in two aspects: First, the adopted multi-modal fusion strategies are built upon well-aligned input images, which are vulnerable to spatial misalignment between modalities (caused by respiratory motions, different scanning parameters, registration errors, etc). Second, the performance of known methods remains subject to the uncertainty of segmentation, which is particularly acute in tumor boundary regions. To tackle these issues, in this paper, we propose a novel multi-modal tumor segmentation method with deformable feature fusion and uncertain region refinement. Concretely, we introduce a deformable aggregation module, which integrates feature alignment and feature aggregation in an ensemble, to reduce inter-modality misalignment and make full use of cross-modal information. Moreover, we devise an uncertain region inpainting module to refine uncertain pixels using neighboring discriminative features. Experiments on two clinical multi-modal tumor datasets demonstrate that our method achieves promising tumor segmentation results and outperforms state-of-the-art methods.
Yue Zhang 0042, Chengtao Peng, Ruofeng Tong 0001, Lanfen Lin, Yen-Wei Chen 0001, Qingqing Chen 0001, Hongjie Hu, Shaohua Kevin Zhou
IEEE Trans. Medical Imaging3
2022 Mixed Transformer U-Net for Medical Image Segmentation
abstract
Though U-Net has achieved tremendous success in medical image segmentation tasks, it lacks the ability to explicitly model long-range dependencies. Therefore, Vision Transformers have emerged as alternative segmentation structures recently, for their innate ability of capturing long-range correlations through Self-Attention (SA). However, Transformers usually rely on large-scale pre-training and have high computational complexity. Furthermore, SA can only model self-affinities within a single sample, ignoring the potential correlations of the overall dataset. To address these problems, we propose a novel Transformer module named Mixed Transformer Module (MTM) for simultaneous inter- and intra- affinities learning. MTM first calculates self-affinities efficiently through our well-designed Local-Global Gaussian-Weighted Self-Attention (LGG-SA). Then, it mines inter-connections between data samples through External Attention (EA). By using MTM, we construct a U-shaped model named Mixed Transformer U-Net (MT-UNet) for accurate medical image segmentation. We test our method on two different public datasets, and the experimental results show that the proposed method achieves better performance over other state-of-the-art methods. The code is available at: https://github.com/Dootmaan/MT-UNet.
Hongyi Wang 0002, Shiao Xie, Lanfen Lin, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001
ICASSP7
2022 ScaleFormer: Revisiting the Transformer-based Backbones from a Scale-wise Perspective for Medical Image Segmentation
abstract
Recently, a variety of vision transformers have been developed as their capability of modeling long-range dependency. In current transformer-based backbones for medical image segmentation, convolutional layers were replaced with pure transformers, or transformers were added to the deepest encoder to learn global context. However, there are mainly two challenges in a scale-wise perspective: (1) intra-scale problem: the existing methods lacked in extracting local-global cues in each scale, which may impact the signal propagation of small objects; (2) inter-scale problem: the existing methods failed to explore distinctive information from multiple scales, which may hinder the representation learning from objects with widely variable size, shape and location. To address these limitations, we propose a novel backbone, namely ScaleFormer, with two appealing designs: (1) A scale-wise intra-scale transformer is designed to couple the CNN-based local features with the transformer-based global cues in each scale, where the row-wise and column-wise global dependencies can be extracted by a lightweight Dual-Axis MSA. (2) A simple and effective spatial-aware inter-scale transformer is designed to interact among consensual regions in multiple scales, which can highlight the cross-scale dependency and resolve the complex scale variations. Experimental results on different benchmarks demonstrate that our Scale-Former outperforms the current state-of-the-art methods. The code is publicly available at: https://github.com/ZJUGiveLab/ScaleFormer.
Huimin Huang 0002, Shiao Xie, Lanfen Lin, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001
IJCAI7
2022 Reconstructing Recognizable 3D Face Shapes based on 3D Morphable Models
abstract
Abstract Many recent works have reconstructed distinctive 3D face shapes by aggregating shape parameters of the same identity and separating those of different people based on parametric models (e.g. 3D morphable models (3DMMs)). However, despite the high accuracy in the face recognition task using these shape parameters, the visual discrimination of face shapes reconstructed from those parameters remains unsatisfactory. Previous works have not answered the following research question: Do discriminative shape parameters guarantee visual discrimination in represented 3D face shapes? This paper analyses the relationship between shape parameters and reconstructed shape geometry, and proposes a novel shape identity‐aware regularization (SIR) loss for shape parameters, aiming at increasing discriminability in both the shape parameter and shape geometry domains. Moreover, to cope with the lack of training data containing both landmark and identity annotations, we propose a network structure and an associated training strategy to leverage mixed data containing either identity or landmark labels. In addition, since face recognition accuracy does not mean the recognizability of reconstructed face shapes from the shape parameters, we propose the SIR metric to measure the discriminability of face shapes. We compare our method with existing methods in terms of the reconstruction error, visual discriminability, and face recognition accuracy of the shape parameters and SIR metric. Experimental results show that our method outperforms the state‐of‐the‐art methods. The code will be released at https://github.com/a686432/SIR .
Diqiong Jiang, Yiwei Jin, Yukun Lai, Risheng Deng, Ruofeng Tong 0001, Min Tang 0001
Comput. Graph. Forum6
2022 N-Cloth: Predicting 3D Cloth Deformation with Mesh-Based Networks
abstract
Abstract We present a novel mesh‐based learning approach (N‐Cloth) for plausible 3D cloth deformation prediction. Our approach is general and can handle cloth or obstacles represented by triangle meshes with arbitrary topologies. We use graph convolution to transform the cloth and object meshes into a latent space to reduce the non‐linearity in the mesh space. Our network can predict the target 3D cloth mesh deformation based on the initial state of the cloth mesh template and the target obstacle mesh. Our approach can handle complex cloth meshes with up to 100 K triangles and scenes with various objects corresponding to SMPL humans, non‐SMPL humans or rigid bodies. In practice, our approach can be used to generate plausible cloth simulation at 30 – 45 fps on an NVIDIA GeForce RTX 3090 GPU. We highlight its benefits over prior learning‐based methods and physically‐based cloth simulators.
Yudi Li, Min Tang 0001, Yun Yang 0001, Zi Huang, Ruofeng Tong 0001, Shuangcai Yang, Dinesh Manocha
Comput. Graph. Forum5
2022 3D corrective nose reconstruction from a single image
abstract
There is a steadily growing range of applications that can benefit from facial reconstruction techniques, leading to an increasing demand for reconstruction of high-quality 3D face models. While it is an important expressive part of the human face, the nose has received less attention than other expressive regions in the face reconstruction literature. When applying existing reconstruction methods to facial images, the reconstructed nose models are often inconsistent with the desired shape and expression. In this paper, we propose a coarse-to-fine 3D nose reconstruction and correction pipeline to build a nose model from a single image, where 3D and 2D nose curve correspondences are adaptively updated and refined. We first correct the reconstruction result coarsely using constraints of 3D-2D sparse landmark correspondences, and then heuristically update a dense 3D-2D curve correspondence based on the coarsely corrected result. A final refinement step is performed to correct the shape based on the updated 3D-2D dense curve constraints. Experimental results show the advantages of our method for 3D nose reconstruction over existing methods.
Yanlong Tang, Yun Zhang 0024, Xiaoguang Han 0001, Yukun Lai, Ruofeng Tong 0001
Comput. Vis. Media6
2022 Attention-based cross-layer domain alignment for unsupervised domain adaptation
Junkun Yuan, Yen-Wei Chen 0001, Ruofeng Tong 0001, Lanfen Lin
Neurocomputing4
2022 BADF: Bounding Volume Hierarchies Centric Adaptive Distance Field Computation for Deformable Objects on GPUs
Xiao-Rui Chen, Min Tang 0001, Dinesh Manocha, Ruofeng Tong 0001
J. Comput. Sci. Technol.5
2022 Source-enhanced prototypical alignment for single image 3D model retrieval
abstract
Abstract Single image 3D model retrieval has attracted a lot of attentions with the convenience of organizing large‐scale unlabeled 3D models. Existing methods transfer the knowledge from well‐annotated 2D images (i.e., source domain) to unlabeled 3D models (i.e., target domain) to improve the discriminability of 3D models and align the feature distributions of 2D images and 3D models. However, during the alignment, the feature learning target of improving the discriminability of 3D models sometimes confuses the boundaries between 2D image categories, where prior methods ignore keeping the discriminability of 2D images. Motivated by this observation, we propose a source‐enhanced prototypical alignment framework to first remain the discriminability of 2D images and then guide the category‐level cross‐domain alignment with better image representations. Specifically, a novel separation and compactness loss is proposed for images to separate the samples from different categories and compact the samples within the same category. Then we perform prototypical alignment to make 2D image features assist in the discriminative feature learning for 3D models. We evaluate the proposed method on the commonly used cross‐domain 3D model retrieval benchmarks, namely MI3DOR and MI3DOR‐2, and the results demonstrate the effectiveness of the proposed method.
Dan Song 0006, Chumeng Zhang, Xuanya Li, Ruofeng Tong 0001
Comput. Animat. Virtual Worlds5
2022 High-fidelity 3D face reconstruction with multi-scale details
Yiwei Jin, Diqiong Jiang, Ruofeng Tong 0001
Pattern Recognit. Lett.4
2022 Mutual Information-Based Graph Co-Attention Networks for Multimodal Prior-Guided Magnetic Resonance Imaging Segmentation
abstract
Multimodal magnetic resonance imaging (MRI) provides complementary information about targets, and the segmentation of multimodal MRI is widely used as an essential preprocessing step for initial diagnosis, stage differentiation, and post-treatment efficacy evaluation in clinical situations. For the main modality or each of the modalities, it is important to enhance the visual information by modeling the connection and effectively fusing the features among them. However, the existing methods for multimodal segmentation have a drawback; they coincidentally drop information of individual modality during the fusion process. Recently, graph learning-based methods have been applied in segmentation, and these methods have achieved considerable improvements by modeling the relationships across feature regions and reasoning using global information. In this paper, we propose a graph learning-based approach to efficiently extract modality-specific features and establish regional correspondence effectively among all modalities. In detail, after projecting features into a graph domain and employing graph convolution to propagate information across all regions for learning global modality-specific features, we propose a mutual information-based graph co-attention module to learn the weight coefficients of one bipartite graph constructed by the fully connected graphs having different modalities in the graph domain and by selectively fusing the node features. Based on the deformation diagram between the spatial-graph space and our proposed graph co-attention module, we present a multimodal prior-guided segmentation framework, which uses two strategies for two clinical situations:Modality-Specific Learning StrategyandCo-Modality Learning Strategy. Besides, the improvedCo-Modality Learning Strategyis used with trainable weights in the multi-task loss for the optimization of the proposed framework. We validated our proposed modules and frameworks on two multimodal MRI datasets: our private liver lesion dataset and a public prostate zone dataset. Our experimental results on both datasets prove the superiority of our proposed approaches.
Shaocong Mo, Lanfen Lin, Ruofeng Tong 0001, Qingqing Chen 0001, Fang Wang 0030, Hongjie Hu, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 MTL-ABS3Net: Atlas-Based Semi-Supervised Organ Segmentation Network With Multi-Task Learning for Medical Images
abstract
Organ segmentation is one of the most important step for various medical image analysis tasks. Recently, semi-supervised learning (SSL) has attracted much attentions by reducing labeling cost. However, most of the existing SSLs neglected the prior shape and position information specialized in the medical images, leading to unsatisfactory localization and non-smooth of objects. In this paper, we propose a novel atlas-based semi-supervised segmentation network with multi-task learning for medical organs, named MTL-ABS3Net, which incorporates the anatomical priors and makes full use of unlabeled data in a self-training and multi-task learning manner. The MTL-ABS3Net consists of two components: an Atlas-Based Semi-Supervised Segmentation Network (ABS3Net) and Reconstruction-Assisted Module (RAM). Specifically, the ABS3Net improves the existing SSLs by utilizing atlas prior, which generates credible pseudo labels in a self-training manner; while the RAM further assists the segmentation network by capturing the anatomical structures from the original images in a multi-task learning manner. Better reconstruction quality is achieved by using MS-SSIM loss function, which further improves the segmentation accuracy. Experimental results from the liver and spleen datasets demonstrated that the performance of our method was significantly improved compared to existing state-of-the-art methods.
Huimin Huang 0002, Qingqing Chen 0001, Lanfen Lin, Qiaowei Zhang, Yutaro Iwamoto, Xianhua Han, Akira Furukawa, Shuzo Kanasaki, Yen-Wei Chen 0001, Ruofeng Tong 0001, Hongjie Hu
IEEE J. Biomed. Health Informatics11
2022 DeepRecS: From RECIST Diameters to Precise Liver Tumor Segmentation
abstract
Liver tumor segmentation (LiTS) is of primary importance in diagnosis and treatment of hepatocellular carcinoma. Known automated LiTS methods could not yield satisfactory results for clinical use since they were hard to model flexible tumor shapes and locations. In clinical practice, radiologists usually estimate tumor shape and size by a Response Evaluation Criteria in Solid Tumor (RECIST) mark. Inspired by this, in this paper, we explore a deep learning (DL) based interactive LiTS method, which incorporates guidance from user-provided RECIST marks. Our method takes a three-step framework to predict liver tumor boundaries. Under this architecture, we develop a RECIST mark propagation network (RMP-Net) to estimate RECIST-like marks in off-RECIST slices. We also devise a context-guided boundary-sensitive network (CGBS-Net) to distill tumors' contextual and boundary information from corresponding RECIST(-like) marks, and then predict tumor maps. To further refine the segmentation results, we process the tumor maps using a 3D conditional random field (CRF) algorithm and a morphology hole-filling operation. Verified on two clinical contrast-enhanced abdomen computed tomography (CT) image datasets, our proposed approach can produce promising segmentation results, and outperforms the state-of-the-art interactive segmentation methods.
Yue Zhang 0042, Chengtao Peng, Liying Peng, Lanfen Lin, Ruofeng Tong 0001, Zhiyi Peng, Xiongwei Mao, Hongjie Hu, Yen-Wei Chen 0001, Jingsong Li 0001
IEEE J. Biomed. Health Informatics6
2021 Toward Realistic Virtual Try-on Through Landmark Guided Shape Matching
abstract
Image-based virtual try-on aims to synthesize the customer image with an in-shop clothes image to acquire seamless and natural try-on results, which have attracted increasing attentions. The main procedures of image-based virtual try-on usually consist of clothes image generation and try-on image synthesis, whereas prior arts cannot guarantee satisfying clothes results when facing large geometric changes and complex clothes patterns, which further deteriorates the afterwards try-on results. To address this issue, we propose a novel virtual try-on network based on landmark-guided shape matching (LM-VTON). Specifically, the clothes image generation progressively learns the warped clothes and refined clothes in an end-to-end manner, where we introduce a landmark-based constraint in Thin-Plate Spline (TPS) warping to inject finer deformation constraints around the clothes. The try-on process synthesizes the warped clothes with personal characteristics via a semantic indicator. Qualitative and quantitative experiments on two public datasets validate the superiority of the proposed method, especially for challenging cases such as large geometric changes and complex clothes patterns. Code will be available at https://github.com/lgqfhwy/LM-VTON.
Dan Song 0006, Ruofeng Tong 0001, Min Tang 0001
AAAI3
2021 Graph-Based Pyramid Global Context Reasoning With a Saliency- Aware Projection for Covid-19 Lung Infections Segmentation
abstract
Coronavirus Disease 2019 (COVID-19) has rapidly spread in 2020, emerging a mass of studies for lung infection segmentation from CT images. Though many methods have been proposed for this issue, it is a challenging task because of infections of various size appearing in different lobe zones. To tackle these issues, we propose a Graph-based Pyramid Global Context Reasoning (Graph-PGCR) module, which is capable of modeling long-range dependencies among disjoint infections as well as adapt size variation. We first incorporate graph convolution to exploit long-term contextual information from multiple lobe zones. Different from previous average pooling or maximum object probability, we propose a saliency-aware projection mechanism to pick up infection-related pixels as a set of graph nodes. After graph reasoning, the relation-aware features are reversed back to the original coordinate space for the down-stream tasks. We further construct multiple graphs with different sampling rates to handle the size variation problem. To this end, distinct multi-scale long-range contextual patterns can be captured. Our Graph- PGCR module is plug-and-play, which can be integrated into any architecture to improve its performance. Experiments demonstrated that the proposed method consistently boost the performance of state-of-the-art backbone architectures on both of public and our private COVID-19 datasets.
Huimin Huang 0002, Lanfen Lin, Xiongwei Mao, Xiaohan Qian, Zhiyi Peng, Jianying Zhou 0006, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001
ICASSP12
2021 Graph-BAS3Net: Boundary-Aware Semi-Supervised Segmentation Network with Bilateral Graph Convolution
abstract
Semi-supervised learning (SSL) algorithms have attracted much attentions in medical image segmentation by leveraging unlabeled data, which challenge in acquiring massive pixel-wise annotated samples. However, most of the existing SSLs neglected the geometric shape constraint in object, leading to unsatisfactory boundary and non-smooth of object. In this paper, we propose a novel boundary-aware semi-supervised medical image segmentation network, named Graph-BAS3Net, which incorporates the boundary information and learns duality constraints between semantics and geometrics in the graph domain. Specifically, the proposed method consists of two components: a multi-task learning framework BAS3Net and a graph-based cross-task module BGCM. The BAS3Net improves the existing GAN-based SSL by adding a boundary detection task, which encodes richer features of object shape and surface. Moreover, the BGCM further explores the co-occurrence relations between the semantics segmentation and boundary detection task, so that the network learns stronger semantic and geometric correspondences from both labeled and unlabeled data. Experimental results on the LiTS dataset and COVID-19 dataset confirm that our proposed Graph-BAS3Net outperforms the state-of-the-art methods in semi-supervised segmentation task.
Huimin Huang 0002, Lanfen Lin, Yue Zhang 0042, Xiongwei Mao, Xiaohan Qian, Zhiyi Peng, Jianying Zhou 0006, Yen-Wei Chen 0001, Ruofeng Tong 0001
ICCV11
2021 3D Graph-S2Net: Shape-Aware Self-ensembling Network for Semi-supervised Segmentation with Bilateral Graph Convolution
Huimin Huang 0002, Lanfen Lin, Hongjie Hu, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001
MICCAI (2)8
2021 Patch-Free 3D Medical Image Segmentation Driven by Super-Resolution Technique and Self-Supervised Guidance
Hongyi Wang 0002, Lanfen Lin, Hongjie Hu, Qingqing Chen 0001, Yinhao Li 0002, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001
MICCAI (1)9
2021 Multi-phase Liver Tumor Segmentation with Spatial Aggregation and Uncertain Region Inpainting
Yue Zhang 0042, Chengtao Peng, Liying Peng, Huimin Huang 0002, Ruofeng Tong 0001, Lanfen Lin, Jingsong Li 0001, Yen-Wei Chen 0001, Qingqing Chen 0001, Hongjie Hu, Zhiyi Peng
MICCAI (1)5
2021 M-DFNet: Multi-phase Discriminative Feature Network for Retrieval of Focal Liver Lesions
abstract
Content based medical image retrieval (CBMIR) plays a great role in computer aided diagnosis for assisting radiologists to detect and characterize focal liver lesions (FLLs). Deep learning has gained exciting performance on CBMIR. While the features generated by deep learning models trained using softmax loss are always separable but not discriminative enough, which is insufficient for retrieval task. In this paper, we propose a multi-phase discriminative feature network (M-DFNet) with a DeepExtracter and a feature refine module (FRModule) to learn discriminative and separable features under a joint supervision of center loss and softmax loss. The hybrid loss enables to minimize intra-class variations and enlarge inter-class differences as much as possible. The FRModule is proposed to recalibrate the deep features based on the learned class centers to tackle the complex imaging manifestations of FLLs and further enhance both the feature discrimination and generalization. Multi-phase computed tomography (CT) images contain pivotal information for diagnosis of FLLs. Thus the M-DFNet is designed to cope with multi-phase information and we explore an appropriate and effective method for multi-phase feature integration on limited data. Experimental results clearly demonstrate strong performance superiority by our proposed method.
Jing Liu 0041, Lanfen Lin, Hongjie Hu, Ruofeng Tong 0001, Jingsong Li 0001, Yen-Wei Chen 0001
ICMR5
2021 Accurate and fast mitotic detection using an anchor-free method based on full-scale connection with recurrent deep layer aggregation in 4D microscopy images
abstract
BACKGROUND: To effectively detect and investigate various cell-related diseases, it is essential to understand cell behaviour. The ability to detection mitotic cells is a fundamental step in diagnosing cell-related diseases. Convolutional neural networks (CNNs) have been successfully applied to object detection tasks, however, when applied to mitotic cell detection, most existing methods generate high false-positive rates due to the complex characteristics that differentiate normal cells from mitotic cells. Cell size and orientation variations in each stage make detecting mitotic cells difficult in 2D approaches. Therefore, effective extraction of the spatial and temporal features from mitotic data is an important and challenging task. The computational time required for detection is another major concern for mitotic detection in 4D microscopic images. RESULTS: In this paper, we propose a backbone feature extraction network named full scale connected recurrent deep layer aggregation (RDLA++) for anchor-free mitotic detection. We utilize a 2.5D method that includes 3D spatial information extracted from several 2D images from neighbouring slices that form a multi-stream input. CONCLUSIONS: Our proposed technique addresses the scale variation problem and can efficiently extract spatial and temporal features from 4D microscopic images, resulting in improved detection accuracy and reduced computation time compared with those of other state-of-the-art methods.
Titinunt Kitrungrotsakul, Yutaro Iwamoto, Satoko Takemoto, Hideo Yokota, Sari Ipponjima, Tomomi Nemoto, Lanfen Lin, Ruofeng Tong 0001, Jingsong Li 0001, Yen-Wei Chen 0001
BMC Bioinform.8
2021 VolumeNet: A Lightweight Parallel Network for Super-Resolution of MR and CT Volumetric Data
abstract
Deep learning-based super-resolution (SR) techniques have generally achieved excellent performance in the computer vision field. Recently, it has been proven that three-dimensional (3D) SR for medical volumetric data delivers better visual results than conventional two-dimensional (2D) processing. However, deepening and widening 3D networks increases training difficulty significantly due to the large number of parameters and small number of training samples. Thus, we propose a 3D convolutional neural network (CNN) for SR of magnetic resonance (MR) and computer tomography (CT) volumetric data called ParallelNet using parallel connections. We construct a parallel connection structure based on the group convolution and feature aggregation to build a 3D CNN that is as wide as possible with a few parameters. As a result, the model thoroughly learns more feature maps with larger receptive fields. In addition, to further improve accuracy, we present an efficient version of ParallelNet (called VolumeNet), which reduces the number of parameters and deepens ParallelNet using a proposed lightweight building block module called the Queue module. Unlike most lightweight CNNs based on depthwise convolutions, the Queue module is primarily constructed using separable 2D cross-channel convolutions. As a result, the number of network parameters and computational complexity can be reduced significantly while maintaining accuracy due to full channel fusion. Experimental results demonstrate that the proposed VolumeNet significantly reduces the number of model parameters and achieves high precision results compared to state-of-the-art methods in tasks of brain MR image SR, abdomen CT image SR, and reconstruction of super-resolution 7T-like images from their 3T counterparts.
Yinhao Li 0002, Yutaro Iwamoto, Lanfen Lin, Rui Xu 0002, Ruofeng Tong 0001, Yen-Wei Chen 0001
IEEE Trans. Image Process.5
2021 Attention-RefNet: Interactive Attention Refinement Network for Infected Area Segmentation of COVID-19
abstract
COVID-19 pneumonia is a disease that causes an existential health crisis in many people by directly affecting and damaging lung cells. The segmentation of infected areas from computed tomography (CT) images can be used to assist and provide useful information for COVID-19 diagnosis. Although several deep learning-based segmentation methods have been proposed for COVID-19 segmentation and have achieved state-of-the-art results, the segmentation accuracy is still not high enough (approximately 85%) due to the variations of COVID-19 infected areas (such as shape and size variations) and the similarities between COVID-19 and non-COVID-infected areas. To improve the segmentation accuracy of COVID-19 infected areas, we propose an interactive attention refinement network (Attention RefNet). The interactive attention refinement network can be connected with any segmentation network and trained with the segmentation network in an end-to-end fashion. We propose a skip connection attention module to improve the important features in both segmentation and refinement networks and a seed point module to enhance the important seeds (positions) for interactive refinement. The effectiveness of the proposed method was demonstrated on public datasets (COVID-19CTSeg and MICCAI) and our private multicenter dataset. The segmentation accuracy was improved to more than 90%. We also confirmed the generalizability of the proposed network on our multicenter dataset. The proposed method can still achieve high segmentation accuracy.
Titinunt Kitrungrotsakul, Qingqing Chen 0001, Huitao Wu, Yutaro Iwamoto, Hongjie Hu, Wenchao Zhu, Fangyi Xu, Lanfen Lin, Ruofeng Tong 0001, Jingsong Li 0001, Yen-Wei Chen 0001
IEEE J. Biomed. Health Informatics11
2021 Medical Image Segmentation With Deep Atlas Prior
abstract
Organ segmentation from medical images is one of the most important pre-processing steps in computer-aided diagnosis, but it is a challenging task because of limited annotated data, low-contrast and non-homogenous textures. Compared with natural images, organs in the medical images have obvious anatomical prior knowledge (e.g., organ shape and position), which can be used to improve the segmentation accuracy. In this paper, we propose a novel segmentation framework which integrates the medical image anatomical prior through loss into the deep learning models. The proposed prior loss function is based on probabilistic atlas, which is called as deep atlas prior (DAP). It includes prior location and shape information of organs, which are important prior information for accurate organ segmentation. Further, we combine the proposed deep atlas prior loss with the conventional likelihood losses such as Dice loss and focal loss into an adaptive Bayesian loss in a Bayesian framework, which consists of a prior and a likelihood. The adaptive Bayesian loss dynamically adjusts the ratio of the DAP loss and the likelihood loss in the training epoch for better learning. The proposed loss function is universal and can be combined with a wide variety of existing deep segmentation models to further enhance their performance. We verify the significance of our proposed framework with some state-of-the-art models, including fully-supervised and semi-supervised segmentation models on a public dataset (ISBI LiTS 2017 Challenge) for liver segmentation and a private dataset for spleen segmentation.
Huimin Huang 0002, Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001
IEEE Trans. Medical Imaging11
2020 UNet 3+: A Full-Scale Connected UNet for Medical Image Segmentation
abstract
Recently, a growing interest has been seen in deep learning-based semantic segmentation. UNet, which is one of deep learning networks with an encoder-decoder architecture, is widely used in medical image segmentation. Combining multi-scale features is one of important factors for accurate segmentation. UNet++ was developed as a modified Unet by designing an architecture with nested and dense skip connections. However, it does not explore sufficient information from full scales and there is still a large room for improvement. In this paper, we propose a novel UNet 3+, which takes advantage of full-scale skip connections and deep supervisions. The full-scale skip connections incorporate low-level details with high-level semantics from feature maps in different scales; while the deep supervision learns hierarchical representations from the full-scale aggregated feature maps. The proposed method is especially benefiting for organs that appear at varying scales. In addition to accuracy improvements, the proposed UNet 3+ can reduce the network parameters to improve the computation efficiency. We further propose a hybrid loss function and devise a classification-guided module to enhance the organ boundary and reduce the over-segmentation in a non-organ image, yielding more accurate segmentation results. The effectiveness of the proposed method is demonstrated on two datasets. The code is available at: github.com/ZJUGiveLab/UNet-Version.
Huimin Huang 0002, Lanfen Lin, Ruofeng Tong 0001, Hongjie Hu, Qiaowei Zhang, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Jian Wu 0001
ICASSP3
2020 Multimodal Priors Guided Segmentation of Liver Lesions in MRI Using Mutual Information Based Graph Co-Attention Networks
Shaocong Mo, Lanfen Lin, Ruofeng Tong 0001, Qingqing Chen 0001, Fang Wang 0030, Hongjie Hu, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001
MICCAI (4)4
2020 P-cloth: interactive complex cloth simulation on multi-GPU systems using dynamic matrix assembly and pipelined implicit integrators
abstract
We present a novel parallel algorithm for cloth simulation that exploits multiple GPUs for fast computation and the handling of very high resolution meshes. To accelerate implicit integration, we describe new parallel algorithms for sparse matrix-vector multiplication (SpMV) and for dynamic matrix assembly on a multi-GPU workstation. Our algorithms use a novel work queue generation scheme for a fat-tree GPU interconnect topology. Furthermore, we present a novel collision handling scheme that uses spatial hashing for discrete and continuous collision detection along with a non-linear impact zone solver. Our parallel schemes can distribute the computation and storage overhead among multiple GPUs and enable us to perform almost interactive simulation on complex cloth meshes, which can hardly be handled on a single GPU due to memory limitations. We have evaluated the performance with two multi-GPU workstations (with 4 and 8 GPUs, respectively) on cloth meshes with 0.5 -- 1.65 M triangles. Our approach can reliably handle the collisions and generate vivid wrinkles and folds at 2 -- 5 fps, which is significantly faster than prior cloth simulation systems. We observe almost linear speedups with respect to the number of GPUs.
Min Tang 0001, Ruofeng Tong 0001, Jieyi Zhao, Dinesh Manocha
ACM Trans. Graph.3
2019 A Dual-Attention Dilated Residual Network for Liver Lesion Classification and Localization on CT Images
abstract
Automatic liver lesion classification on computed tomography images is of great importance to early cancer diagnosis and remains a challenging task. State-of-the-art liver lesion classification algorithms are currently based on manually selected regions of interest (ROIs) or automatically detected ROIs. However, liver lesions usually vary in size and shape, which makes the ROI selection process labor-intensive and also poses an obstacle to automatic lesion detection. In this paper, we propose a dual-attention dilated residual network (DADRN) as a potential solution to lesion classification task without manual ROI selection or automatic lesion detection. We incorporated a novel dual-attention module in order to capture the non-local feature dependencies and help the deep neural network focus on the lesion area by enlarging the difference between the lesion area and nonlesion area. To the best of our knowledge, we are the first to employ the self-attention mechanism to address liver lesion classification task. In addition, the well-trained DADRN can be used for weakly-supervised lesion localization without any architectural change or retraining. Experiment results show that DADRN could achieve a lesion classification accuracy comparable to that of the state-of-the-art ROI-based method and outperformed state-of-the-art attention-based approaches in both liver lesion classification and localization tasks.
Xiao Chen 0016, Jian Wu 0001, Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001
ICIP10
2019 Multi-Stream Scale-Insensitive Convolutional and Recurrent Neural Networks for Liver Tumor Detection in Dynamic Ct Images
abstract
Convolutional neural networks (CNNs) have achieved great success in numerous challenging vision tasks, and have great potential for object detection in natural images. Compared with the natural images, medical images exhibit some unique characteristics. Therefore, substantial challenges still remain in this field. The first challenge is to develop a method for effectively distilling enhancement patterns from the dynamic CT images. Moreover, since tumor sizes vary greatly and small lesions are important for early liver tumor detection, lesion detection with a widely variable scale is another challenge. In this paper, we propose a multi-stream scale-insensitive convolutional and recurrent neural network (MSCR) for liver tumor detection. Specifically, we propose the use of grouped convolutional long short-term memory (GCLSTM) to extract enhancement patterns, which is developed as a plug-and-play module. Experiments show that the MSCR framework exhibits superior performance over state-of-the-art approaches, achieving an average precision of 77.06% for detection of focal liver lesions. We have released the code of MSCR in1.
Ruofeng Tong 0001, Jian Wu 0001, Lanfen Lin, Xiao Chen 0016, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001
ICIP2
2019 Semi-supervised Segmentation of Liver Using Adversarial Learning with Deep Atlas Prior
Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001, Jian Wu 0001
MICCAI (6)9
2019 A three-stage real-time detector for traffic signs in large panoramas
abstract
Traffic sign detection is one of the key components in autonomous driving. Advanced autonomous vehicles armed with high quality sensors capture high definition images for further analysis. Detecting traffic signs, moving vehicles, and lanes is important for localization and decision making. Traffic signs, especially those that are far from the camera, are small, and so are challenging to traditional object detection methods. In this work, in order to reduce computational cost and improve detection performance, we split the large input images into small blocks and then recognize traffic signs in the blocks using another detection module. Therefore, this paper proposes a three-stage traffic sign detector, which connects a BlockNet with an RPN–RCNN detection network. BlockNet, which is composed of a set of CNN layers, is capable of performing block-level foreground detection, making inferences in less than 1 ms. Then, the RPN–RCNN two-stage detector is used to identify traffic sign objects in each block; it is trained on a derived dataset named TT100KPatch. Experiments show that our framework can achieve both state-of-the-art accuracy and recall; its fastest detection speed is 102 fps.
Ruochen Fan, Sharon X. Huang, Zhe Zhu, Ruofeng Tong 0001
Comput. Vis. Media5
2019 Illumination-aware faster R-CNN for robust multispectral pedestrian detection
Dan Song 0006, Ruofeng Tong 0001, Min Tang 0001
Pattern Recognit.3
2019 Expressive facial style transfer for personalized memes mimic
Yanlong Tang, Xiaoguang Han 0001, Yue Li 0049, Liqian Ma, Ruofeng Tong 0001
Vis. Comput.5
2018 Multispectral Pedestrian Detection via Simultaneous Detection and Segmentation
Dan Song 0006, Ruofeng Tong 0001, Min Tang 0001
BMVC3
2018 Accurate self-collision detection using enhanced dual-cone method
Min Tang 0001, Zhendong Wang 0001, Ruofeng Tong 0001
Comput. Graph.4
2018 Efficient BVH-based Collision Detection Scheme with Ordering and Restructuring
abstract
Abstract Bounding volume hierarchy (BVH) has been widely adopted as the acceleration structure in broad‐phase collision detection. Previous state‐of‐the‐art BVH‐based collision detection approaches exploited the spatio‐temporal coherence of simulations by maintaining a bounding volume test tree (BVTT) front. A major drawback of these algorithms is that large deformations in the scenes decrease culling efficiency and slow down collision queries. Moreover, for front‐based methods, the inefficient caching on GPU caused by the arbitrary layout of BVH and BVTT front nodes becomes a critical performance issue. We present a fast and robust BVH‐based collision detection scheme on GPU that addresses the above problems by ordering and restructuring BVHs and BVTT fronts. Our techniques are based on the use of histogram sort and an auxiliary structure BVTT front log, through which we analyze the dynamic status of BVTT front and BVH quality. Our approach efficiently handles inter‐ and intra‐object collisions and performs especially well in simulations where there is considerable spatio‐temporal coherence. The benchmark results demonstrate that our approach is significantly faster than the previous BVH‐based method, and also outperforms other state‐of‐the‐art spatial subdivision schemes in terms of speed.
Min Tang 0001, Dinesh Manocha, Ruofeng Tong 0001
Comput. Graph. Forum4
2018 I-cloth: incremental collision handling for GPU-based interactive cloth simulation
abstract
We present an incremental collision handling algorithm for GPU-based interactive cloth simulation. Our approach exploits the spatial and temporal coherence between successive iterations of an optimization-based solver for collision response computation. We present an incremental continuous collision detection algorithm that keeps track of deforming vertices and combine it with spatial hashing. We use a non-linear GPU-based impact zone solver to resolve the penetrations. We combine our collision handling algorithm with implicit integration to use large time steps. Our overall algorithm, I-Cloth, can simulate complex cloth deformation with a few hundred thousand vertices at 2 - 8 frames per second on a commodity GPU. We highlight its performance on different benchmarks and observe up to 7 - 10X speedup over prior algorithms.
Min Tang 0001, Zhongyuan Liu, Ruofeng Tong 0001, Dinesh Manocha
ACM Trans. Graph.4
2017 Efficient and Reliable Self-Collision Culling Using Unprojected Normal Cones
abstract
Abstract We present an efficient and accurate algorithm for self‐collision detection in deformable models. Our approach can perform discrete and continuous collision queries on triangulated meshes. We present a simple and linear time algorithm to perform the normal cone test using the unprojected 3D vertices, which reduces to a sequence point‐plane classification tests. Moreover, we present a hierarchical traversal scheme that can significantly reduce the number of normal cone tests and the memory overhead using front‐based normal cone culling. The overall algorithm can reliably detect all (self) collisions in models composed of hundreds of thousands of triangles. We observe considerable performance improvement over prior continuous collision detection algorithms.
Min Tang 0001, Ruofeng Tong 0001, Dinesh Manocha
Comput. Graph. Forum4
2017 Patch-based variational image approximation
Ruofeng Tong 0001
Sci. China Inf. Sci.2
2017 Exploitation of multiplayer interaction and development of virtual puppetry storytelling using gesture control and stereoscopic devices
abstract
Abstract With the rapid development of human–computer interaction technologies, the new media generation demands novel learning experiences with natural interaction and immersive experience. Considering that digital storytelling is a powerful pedagogical tool for young children, in this paper, we design an immersive storytelling environment that allows multiple players to use naturally interactive hand gestures to manipulate virtual puppetry for assisting narration. A set of multimodal interaction techniques is presented for a hybrid user interface that integrates existing 3D visualization and interaction devices including head‐mounted displays and depth motion sensor. In this system, the young players could intuitively use hand gestures to manipulate virtual puppets to perform a story and interact with props in a virtual stereoscopic environment. We have conducted a user experiment with four young children for pedagogical evaluation, as well as system acceptability and interactivity evaluation by postgraduate students. The results show that our framework has great potential to stimulate learning abilities of young children through collaboration tasks. The stereoscopic head‐mounted display outperformed the traditional monoscopic display in a comparison between the two.
Hui Liang 0004, Jian Chang 0001, Shujie Deng, Can Chen 0001, Ruofeng Tong 0001, Jian J. Zhang 0001
Comput. Animat. Virtual Worlds5
2016 3D Body Shapes Estimation from Dressed-Human Silhouettes
abstract
Abstract Estimation of 3D body shapes from dressed‐human photos is an important but challenging problem in virtual fitting. We propose a novel automatic framework to efficiently estimate 3D body shapes under clothes. We construct a database of 3D naked and dressed body pairs, based on which we learn how to predict 3D positions of body landmarks (which further constrain a parametric human body model) automatically according to dressed‐human silhouettes. Critical vertices are selected on 3D registered human bodies as landmarks to represent body shapes, so as to avoid the time‐consuming vertices correspondences finding process for parametric body reconstruction. Our method can estimate 3D body shapes from dressed‐human silhouettes within 4 seconds, while the fastest method reported previously need 1 minute. In addition, our estimation error is within the size tolerance for clothing industry. We dress 6042 naked bodies with 3 sets of common clothes by physically based cloth simulation technique. To the best of our knowledge, We are the first to construct such a database containing 3D naked and dressed body pairs and our database may contribute to the areas of human body shapes estimation and cloth simulation.
Dan Song 0006, Ruofeng Tong 0001, Jian Chang 0001, Xiaosong Yang, Min Tang 0001, Jian J. Zhang 0001
Comput. Graph. Forum2
2016 CAMA: Contact-Aware Matrix Assembly with Unified Collision Handling for GPU-based Cloth Simulation
abstract
Abstract We present a novel GPU‐based approach to robustly and efficiently simulate high‐resolution and complexly layered cloth. The key component of our formulation is a parallelized matrix assembly algorithm that can quickly build a large and sparse matrix in a compressed format and accurately solve linear systems on GPUs. We also present a fast and integrated solution for parallel collision handling, including collision detection and response computations, which utilizes spatio‐temporal coherence. We combine these algorithms as part of a new cloth simulation pipeline that incorporates contact forces into implicit time integration for collision avoidance. The entire pipeline is implemented on GPUs, and we evaluate its performance on complex benchmarks consisting of 100 – 300K triangles. In practice, our system takes a few seconds to simulate one frame of a complex cloth scene, which represents significant speedups over prior CPU and GPU‐based cloth simulation systems.
Min Tang 0001, Huamin Wang 0001, Le Tang, Ruofeng Tong 0001, Dinesh Manocha
Comput. Graph. Forum4
2016 Efficient and robust strain limiting and treatment of simultaneous collisions with semidefinite programming
abstract
We present an efficient and robust method which performs well for both strain limiting and treatment of simultaneous collisions. Our method formulates strain constraints and collision constraints as a serial of linear matrix inequalities (LMIs) and linear polynomial inequalities (LPIs), and solves an optimization problem with standard convex semidefinite programming solvers. When performing strain limiting, our method acts on strain tensors to constrain the singular values of the deformation gradient matrix in a specified interval. Our method can be applied to both triangular surface meshes and tetrahedral volume meshes. Compared with prior strain limiting methods, our method converges much faster and guarantees triangle flipping does not occur when applied to a triangular mesh. When performing treatment of simultaneous collisions, our method eliminates all detected collisions during each iteration, leading to higher efficiency and faster convergence than prior collision treatment methods.
Zhendong Wang 0001, Min Tang 0001, Ruofeng Tong 0001
Comput. Vis. Media4
2016 A Linear Approach for Depth and Colour Camera Calibration Using Hybrid Parameters
Ke-Li Cheng, Xuan Ju, Ruofeng Tong 0001, Min Tang 0001, Jian Chang 0001, Jian J. Zhang 0001
J. Comput. Sci. Technol.3
2016 Image meshing via hierarchical optimization
abstract
Vector graphic, as a kind of geometric representation of raster images, has many advantages, e.g., definition independence and editing facility. A popular way to convert raster images into vector graphics is image meshing, the aim of which is to find a mesh to represent an image as faithfully as possible. For traditional meshing algorithms, the crux of the problem resides mainly in the high non-linearity and non-smoothness of the objective, which makes it difficult to find a desirable optimal solution. To ameliorate this situation, we present a hierarchical optimization algorithm solving the problem from coarser levels to finer ones, providing initialization for each level with its coarser ascent. To further simplify the problem, the original non-convex problem is converted to a linear least squares one, and thus becomes convex, which makes the problem much easier to solve. A dictionary learning framework is used to combine geometry and topology elegantly. Then an alternating scheme is employed to solve both parts. Experiments show that our algorithm runs fast and achieves better results than existing ones for most images.
Ruofeng Tong 0001
Frontiers Inf. Technol. Electron. Eng.2
2016 Parametric Human Body Reconstruction Based on Sparse Key Points
abstract
We propose an automatic parametric human body reconstruction algorithm which can efficiently construct a model using a single Kinect sensor. A user needs to stand still in front of the sensor for a couple of seconds to measure the range data. The user's body shape and pose will then be automatically constructed in several seconds. Traditional methods optimize dense correspondences between range data and meshes. In contrast, our proposed scheme relies on sparse key points for the reconstruction. It employs regression to find the corresponding key points between the scanned range data and some annotated training data. We design two kinds of feature descriptors as well as corresponding regression stages to make the regression robust and accurate. Our scheme follows with dense refinement where a pre-factorization method is applied to improve the computational efficiency. Compared with other methods, our scheme achieves similar reconstruction accuracy but significantly reduces runtime.
Ke-Li Cheng, Ruofeng Tong 0001, Min Tang 0001, Jing-Ye Qian, Michel Sarkis
IEEE Trans. Vis. Comput. Graph.2
2016 Interactive mesh cloning driven by boundary loop
Guiping Qian, Min Tang 0001, Ruofeng Tong 0001, Ruifang Pan
Vis. Comput.3
2016 Depth incorporating with color improves salient object detection
Yan-Long Tang, Ruofeng Tong 0001, Min Tang 0001, Yun Zhang 0024
Vis. Comput.2
2015 Image-Based Hair Pre-processing for Art Creation: A Case Study of Bas-Relief Modelling
abstract
To better capture the shapes as well as the rich dynamics of hair, image based modelling techniques have been developed for reconstructing their 3D geometry and important visual features. Most hair images contain inevitable noises which impair reconstructed hair models. Therefore we propose to pre-process hair images and provide an orientation map of hair strands to enhance the follow-on modelling. To demonstrate the usage of pre-processing techniques, we apply our pre-processing results for bas-relief stylisation and modelling of hair from image inputs. We compare different techniques to estimate hair orientations, adopting four types of filter mechanisms. Our analysis of their performance sheds insight on designing a suitable pre-processing technique for hair reconstruction from images. Several examples of bas-relief creation validate the effectiveness of the proposed approach.
Wenshu Zhang, Jian Chang 0001, Jian J. Zhang 0001, Meili Wang 0001, Ruofeng Tong 0001
IV5
2015 TightCCD: Efficient and Robust Continuous Collision Detection using Tight Error Bounds
abstract
http://gamma.cs.unc.edu/BSC/ We present a realtime and reliable continuous collision detection (CCD) algorithm between triangulated models that exploits the floating point hardware capability of current CPUs and GPUs. Our formulation is based on Bernstein Sign Classification that takes advantage of the geometry properties of Bernstein basis and Bézier curves to perform Boolean collision queries. We derive tight numerical error bounds on the computations and employ those bounds to design an accurate algorithm using finite-precision arithmetic. Compared with prior floatingpoint CCD algorithms, our approach eliminates all the false negatives and 90–95% of the false positives. We integrated our algorithm (TightCCD) with physically-based simulation system and observe speedups in collision queries of 5–15X compared with prior reliable CCD algorithms. Furthermore, we demonstrate its benefits in terms of improving the performance or robustness of cloth simulation systems.
Zhendong Wang 0001, Min Tang 0001, Ruofeng Tong 0001, Dinesh Manocha
Comput. Graph. Forum3
2015 GPU based real-time simulation of massive falling leaves
abstract
As an important autumn feature, scenes with large numbers of falling leaves are common in movies and games. However, it is a challenge for computer graphics to simulate such scenes in an authentic and efficient manner. This paper proposes a GPU based approach for simulating the falling motion of many leaves in real time. Firstly, we use a motionsynthesis based method to analyze the falling motion of the leaves, which enables us to describe complex falling trajectories using low-dimensional features. Secondly, we transmit a primitive-motion trajectory dataset together with the low-dimensional features of the falling leaves to video memory, allowing us to execute the appropriate calculations on the GPU.
Jing-Ye Qian, Ruofeng Tong 0001, Jian Chang 0001, Jian J. Zhang 0001
Comput. Vis. Media3
2015 Stretch-Minimizing Volumetric Parameterization
Guiping Qian, Jieyi Zhao, Jian Chang 0001, Ruofeng Tong 0001, Jian J. Zhang 0001
J. Comput. Sci. Technol.5
2014 Remeshing-assisted Optimization for Locally Injective Mappings
abstract
Abstract Constructing locally injective mappings for 2D triangular meshes is vital in applications such as deformations. In such a highly constrained optimization, the prescribed tessellation may impose strong restriction on the solution. As a consequence, the feasible region may be too small to contain an ideal solution, which leads to problems of slow convergence, poor solution, or even that no solution can be found. We propose to integrate adaptive remeshing into interior point method to solve this issue. We update the vertex positions via a parameter‐free relaxation enhanced geometry optimization, and then use edge‐flip operations to reduce the residual and keep a reasonable condition number for better convergence. For more robustness, when the iteration of interior point method terminates but leaves the positional constraints unsatisfied, we estimate the edges in the current tessellation that block vertices moving based on the convergence information of the optimization, and then split neighboring edges to break the restriction. The results show that our method has better performance than the solely geometric optimization approaches, especially for extreme deformations.
Jin Huang 0001, Ruofeng Tong 0001
Comput. Graph. Forum3
2014 Content-aware texture mapping
Zeyun Shi, Jin Huang 0001, Ruofeng Tong 0001
Graph. Model.5
2013 A GPU-based Streaming Algorithm for High-Resolution Cloth Simulation
abstract
Abstract We present a GPU‐based streaming algorithm to perform high‐resolution and accurate cloth simulation. We map all the components of cloth simulation pipeline, including time integration, collision detection, collision response, and velocity updating to GPU‐based kernels and data structures. Our algorithm perform intra‐object and inter‐object collisions, handles contacts and friction, and is able to accurately simulate folds and wrinkles. We describe the streaming pipeline and address many issues in terms of obtaining high throughput on many‐core GPUs. In practice, our algorithm can perform high‐fidelity simulation on a cloth mesh with 2M triangles using 3GB of GPU memory. We highlight the parallel performance of our algorithm on three different generations of GPUs. On a high‐end NVIDIA Tesla K20c, we observe up to two orders of magnitude performance improvement as compared to a single‐threaded CPU‐based algorithm, and about one order of magnitude improvement over a 16‐core CPU‐based parallel implementation.
Min Tang 0001, Ruofeng Tong 0001, Rahul Narain, Chang Meng, Dinesh Manocha
Comput. Graph. Forum2
2013 Efficient dark channel based image dehazing using quadtrees
Ruofeng Tong 0001
Sci. China Inf. Sci.2
2013 Preface
Shi-Min Hu 0001, Daniel Thalmann, Ruofeng Tong 0001
J. Comput. Sci. Technol.3
2013 Upper Body Human Detection and Segmentation in Low Contrast Video
abstract
In the application of extracting human regions from videos, many existing methods may lose their efficacy when illumination varies or the human remains still. To address this problem, we propose a method in this paper for human region detection and segmentation by constructing a generalized human upper body model. The method mainly consists of two main procedures. First, foreground connected regions are extracted by background subtraction from the current frame and classified through a human upper body model pretrained with a support vector machine to determine whether they are human regions. Second, we assign an energy function to the region contour and apply an energy minimization procedure to evolve the contour when human regions are polluted by background; for example, a change in lighting conditions. After finding the optimal contour, we update the background and repeat the procedures in next frame. This feedback strategy rectifies the mistaken background regions promptly and extracts human regions correctly. Our experimental results demonstrate that the proposed method is robust enough to handle videos of low contrast as well as normal conditions.
Ruofeng Tong 0001, Di Xie, Min Tang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2013 StereoPasting: Interactive Composition in Stereoscopic Images
abstract
We propose "StereoPasting," an efficient method for depth-consistent stereoscopic composition, in which a source 2D image is interactively blended into a target stereoscopic image. As we paint "disparity" on a 2D image, the disparity map of the selected region is gradually produced by edge-aware diffusion, and then blended with that of the target stereoscopic image. By considering constraints of the expected disparities and perspective scaling, the 2D object is warped to generate an image pair, which is then blended into the target image pair to get the composition result. The warping is formulated as an energy minimization, which could be solved in real time. We also present an interactive composition system, in which users can edit the disparity maps of 2D images by strokes, while viewing the composition results instantly. Experiments show that our method is intuitive and efficient for interactive stereoscopic composition. A lot of applications demonstrate the versatility of our method.
Ruofeng Tong 0001, Yun Zhang 0024, Ke-Li Cheng
IEEE Trans. Vis. Comput. Graph.1
2012 Mesh Segmentation for Parallel Decompression on GPU
Jieyi Zhao, Min Tang 0001, Ruofeng Tong 0001
CVM3
2012 GPU accelerated convex hull computation
Min Tang 0001, Jieyi Zhao, Ruofeng Tong 0001, Dinesh Manocha
Comput. Graph.3
2012 Connectivity-Based Segmentation for GPU-Accelerated Mesh Decompression
Jieyi Zhao, Min Tang 0001, Ruofeng Tong 0001
J. Comput. Sci. Technol.3
2012 Fast continuous collision culling with deforming noncollinear filters
abstract
ABSTRACT We present a novel culling algorithm that uses deforming noncollinear filters to improve the performance of continuous collision detection (CCD) algorithms. The underlying idea is to use simple and effective filters, deforming noncollinear filters (NCFs), that reduce the number of false positives between the primitives. These filters are derived from the collinear conditions and can be easily combined with other culling methods. We have tested its performance on several benchmarks. Comparing with previous methods, we can reduce the number of false positives significantly and improve the overall performance of CCD algorithms, especially for simulations with large time steps. Copyright © 2012 John Wiley & Sons, Ltd.
Min Tang 0001, Ruofeng Tong 0001
Comput. Animat. Virtual Worlds3
2012 Continuous penalty forces
abstract
We present a simple algorithm to compute continuous penalty forces to determine collision response between rigid and deformable models bounded by triangle meshes. Our algorithm computes a well-behaved solution in contrast to the traditional stability and robustness problems of penalty methods, induced by force discontinuities. We trace contact features along their deforming trajectories and accumulate penalty forces along the penetration time intervals between the overlapping feature pairs. Moreover, we present a closed-form expression to compute the continuous and smooth collision response. Our method has very small additional overhead compared to previous penalty methods, and can significantly improve the stability and robustness. We highlight its benefits on several benchmarks.
Min Tang 0001, Dinesh Manocha, Miguel A. Otaduy, Ruofeng Tong 0001
ACM Trans. Graph.4
2012 All-hexahedral mesh generation via inside-out advancing front based on harmonic fields
Ruofeng Tong 0001
Vis. Comput.2
2012 Robust super resolution of compressed video
Min Tang 0001, Ruofeng Tong 0001
Vis. Comput.3
2011 Collision-streams: fast GPU-based collision detection for deformable models
abstract
We present a fast GPU-based streaming algorithm to perform collision queries between deformable models. Our approach is based on hierarchical culling and reduces the computation to generating different streams. We present a novel stream registration method to compact the streams and efficiently compute the potentially colliding pairs of primitives. We also use a deferred front tracking method to lower the memory overhead. The overall algorithm has been implemented on different GPUs and we have evaluated its performance on non-rigid and deformable simulations. We highlight our speedups over prior CPU-based and GPU-based algorithms. In practice, our algorithm can perform inter-object and intra-object computations on models composed of hundreds of thousands of triangles in tens of milliseconds.
Min Tang 0001, Dinesh Manocha, Jiang Lin, Ruofeng Tong 0001
SI3D4
2011 Feature-preserving mesh denoising based on vertices classification
Zhe Bian, Ruofeng Tong 0001
Comput. Aided Geom. Des.2
2011 Facial hexahedral mesh transferring by volumetric mapping based on harmonic fields
Ruofeng Tong 0001
Comput. Graph.3
2011 Video Brush: A Novel Interface for Efficient Video Cutout
abstract
Abstract We present Video Brush, a novel interface for interactive video cutout. Inspired by the progressive selection scheme in images, our interface is designed to select video objects by painting on successive frames as the video plays. The video objects are progressively selected by solving the graph‐cut based local optimization according to the strokes drawn by the brush on each painted frame. In order to provide users interactive feedback, we accelerate 3D graph‐cut by efficient graph building and multi‐level banded graph‐cut. Experimental results show that our novel interface is both intuitive and efficient for video cutout.
Ruofeng Tong 0001, Yun Zhang 0024
Comput. Graph. Forum1
2011 VolCCD: Fast continuous collision culling between deforming volume meshes
abstract
We present a novel culling algorithm to perform fast and robust continuous collision detection between deforming volume meshes. This includes a continuous separating axis test that can conservatively check whether two volume meshes overlap during a given time interval. In addition, we present efficient methods to eliminate redundant elementary tests between the features (e.g., vertices, edges, and faces) of volume elements (e.g., tetrahedra, hexahedra, triangular prisms, etc.). Our approach is applicable to various deforming meshes, including those with changing topologies, and efficiently computes the first time of contact. We are able to perform inter-object and intra-object collision queries in models represented with tens of thousands of volume elements at interactive rates on a single CPU core. Moreover, we observe more than an order of magnitude performance improvement over prior methods.
Min Tang 0001, Dinesh Manocha, Sung-Eui Yoon, Jae-Pil Heo, Ruofeng Tong 0001
ACM Trans. Graph.6
2011 Selective image abstraction
Ruofeng Tong 0001, Jinxiang Dong
Vis. Comput.2
2011 Environment-Sensitive cloning in images
Yun Zhang 0024, Ruofeng Tong 0001
Vis. Comput.2
2010 Fast continuous collision detection using deforming non-penetration filters
abstract
We present a novel culling algorithm that uses deforming non-penetration filters to improve the performance of continuous collision detection (CCD) algorithms. The underlying idea is to use a simple and effective filter that reduces both the number of false positives and the elementary tests between the primitives. This filter is derived from the coplanarity condition and can be easily combined with other methods used to accelerate CCD. We have implemented the algorithm and tested its performance on many non-rigid simulations. In practice, we can reduce the number of false positives significantly and improve the overall performance of CCD algorithms by 1.5--8.2x.
Min Tang 0001, Dinesh Manocha, Ruofeng Tong 0001
SI3D3
2010 MCCD: Multi-core collision detection between deformable models using front-based decomposition
Min Tang 0001, Dinesh Manocha, Ruofeng Tong 0001
Graph. Model.3
2010 Making Slide Shows with Zoomquilts
Ruofeng Tong 0001, Jinxiang Dong
J. Comput. Sci. Technol.2
2010 Inhomogeneous volumetric Laplacian deformation for rhinoplasty planning and simulation system
abstract
Abstract This paper presents an intuitive rhinoplasty planning and simulation system, to provide high quality prediction of postoperative appearance, and design patient specific nose prosthesis automatically. The key component is a novel volumetric Laplacian deformation tool inspired by the state‐of‐the‐art differential surface deformation techniques. Working on the volumetric domain and incorporating inhomogeneous material from CT data make the new approach suitable for soft tissue simulation. In particular, the system employs a special sketch contour driving deformation interface, which can provide realistic 3D rhinoplasty simulation with intuitive and straightforward 2D manipulation. When satisfied with the appearance, the change of soft tissue before and after simulation is utilized to generate the individual prosthesis model automatically. Clinical validation using post‐operative CT data demonstrated that the system can provide prediction results of high quality. And the surgeons who used the system confirmed that this planning system is attractive and has potential for daily clinical practice. Copyright © 2010 John Wiley & Sons, Ltd.
Ruofeng Tong 0001, Jian-Ping Geng, Min Tang 0001
Comput. Animat. Virtual Worlds2
2010 Content-aware copying and pasting in images
Ruofeng Tong 0001
Vis. Comput.2
2009 Computer aided design and evaluation of new anatomic fixation system on entire pelvic model
abstract
This paper presented a special computer aided procedure to design a new sacroliliac anatomic bar-plate internal fixation system, and evaluated its biomechanical properties on an accurate patient-specific finite element model of entire pelvis, compared with two conventional internal fixation methods. Based on virtual anatomical measure of 30 digital pelvic models reconstructed from CT, an anatomic plate was designed according to the complicated structure of the outer table of the posterior ilium, and was integrated into the complete fixation system. Then, an ad hoc semi-automatic mesh generator was employed to construct a patient-specific finite element model of whole pelvis, including elaborate sacroiliac joints, important pelvic ligaments, and interpubic disc, as well as position-dependent cortical thickness and trabecular bone elastic modulus. Following, one side of sacroiliac joint related ligaments were deleted to simulate a complete unilateral sacroiliac joint disruption. Then the new anatomic fixation system was integrated to fix the fracture, and two comparing models including iliosacral screw fixation and front reconstruction plate fixation were also generated. Finally, all models were simulated under same loading conditions. The results demonstrated that the mechanical stability of the new anatomic fixation system was superior, with obviously improved stress distribution and little displacement, which implied an effective internal fixation method for potential clinical application.
Ruofeng Tong 0001, Min Tang 0001
Symposium on Solid and Physical Modeling2
2009 Multi-core collision detection between deformable models
abstract
We present a new parallel algorithm for interactive and continuous collision detection between deformable models. Our algorithm performs incremental hierarchical computations between successive frames and parallelizes the computation among multiple cores on current CPUs. The main computations include front building and updating and performing the elementary tests between the triangle primitives. The overall algorithm can perform inter- and intra-object collisions at interactive rates on current commodity processors on models with many tens of thousands of triangles. In practice, the performance of our algorithm almost scales linearly with the number of cores.
Min Tang 0001, Dinesh Manocha, Ruofeng Tong 0001
Symposium on Solid and Physical Modeling3
2009 Gradient field based inhomogeneous volumetric mesh deformation for maxillofacial surgery simulation
Ruofeng Tong 0001, Jinxiang Dong, Fu-dong Zhu
Comput. Graph.2
2009 Boosted cascade of scattered rectangle features for object detection
Weize Zhang, Ruofeng Tong 0001, Jinxiang Dong
Sci. China Ser. F Inf. Sci.2
2007 Assembly Sequence Planning in VM System
abstract
Virtual manufacturing (VM) has been proved in different industrial contexts as an effective technology to ensure competitivity and increase profit margins. This paper presents such a system named MEWS, short for module-based extensible integrated VM system. The system now contains 7 modules and owns good extensibility. One of the most important modules is assembling module. It has significant impact on the performance of the system. The adoption of hierarchical assembly model with discrete patches brings many benefits for the whole system as well as the assembling module itself. And the hybrid assembly sequence planning based on subassembly identification with human-interaction expands the planner's scope of application and enhances its robustness, while restricts human-interaction in hierarchical layers of the assembly model.
Weize Zhang, Ruofeng Tong 0001, Jinxiang Dong
CSCWD2
2006 Synergic Production Process Controlling in Supply Chain Management
abstract
This paper analyses the operational coordination mechanisms between the core enterprise and the partners within a supply chain having private local information. For a make to order production setting, a hybrid model of synergic production scheme based on negotiation is proposed. After analyzing the characteristic of synergic production process, this paper establishes the hierarchy of synergic production process monitoring. In the production process monitoring, four monitoring modes are defined according to the significance of the items in the whole product. A Web-based synergic production management system for supply chain has been implemented with programming in ASP and Ms SQL server. The application shows that the approach proposed in this paper is applicable to practical production in supply chain
Tianyang Dong, Ruofeng Tong 0001, Jinxiang Dong
CSCWD3
2006 A Local Registration Approach of Medical Images with Niche Genetic Algorithm
abstract
This paper proposes a local registration approach of medical images with niche genetic algorithm. In our approach, the deformation function is obtained by interpolating discrete point-landmarks using radial basis function with compact support. The locality of deformation function can be conveniently controlled by distributing the point-landmarks into the desired regions, which especially allows us to deal with local changes in medical images. With niche genetic algorithm, the action scope of each point-landmark is optimized to minimize the sum of squared difference of intensity between two images. Experimental results show that the performance of our image registration approach is very promising in controlling the locality of deformation function
Wen Peng, Ruofeng Tong 0001, Guiping Qian, Jinxiang Dong
CSCWD2
2006 An Efficient Method to Mesh Point Cloud
abstract
This paper presents a new reverse engineering method for creating 3D mesh models, which approximate an unorganized noisy point set without orientation information. The new method computes sample points by the extended moving least squares method in adaptive octree cell. The octree subdivision is decided by weighted covariance matrix. Then the points are connected by intersections of supported spheres. Further, the triangular meshes are refined to remove non-manifold parts and holes. The new algorithm allows us to construct mesh models from very large point set quickly
Guiping Qian, Ruofeng Tong 0001, Wen Peng, Jinxiang Dong
CSCWD2
2006 A Constrained Ant Colony Algorithm for Image Registration
Wen Peng, Ruofeng Tong 0001, Guiping Qian, Jinxiang Dong
ICIC (3)2
2005 Rapidly generate lumbar spine volume mesh
abstract
This paper presents a new method for rapidly generating accurate patient-specific lumbar spine volume mesh for finite element analysis. First, initial isosurface of the vertebral body is extracted from CT volume data. Then, a series of non-parallel "best cross-section planes" are placed semi-automatic ally according to the morphologic characteristic of surface model, forming a "non-regular piecewise subspace". This subspace and the embedded surface model are transformed to a "regular subspace" covered by a regular structure grid. Based on the information from the 2D contours of warped surface model, the "structure grid contours" are generated, from which a surface mesh of high quality is generated. At the same time, these grid nodes inside of the surface mesh are recorded as insertion points, and the tetrahedral volume mesh is created quickly by a boundary constrained Delaunay triangulation procedure in the "regular subspace". Finally a two-step reverse transform procedure is employed to recover the shape feature of the lumbar volume mesh in the original three-dimensional space, achieving a smooth change of element size transition.
Ruofeng Tong 0001, Minke Wang, Jinxiang Dong
CAD/Graphics2
2005 3D whole tooth model from CT volume using thin-plate splines
abstract
As the tooth root has similar bone density with the jaw where it is embedded, its complete boundaries are either missing or at low contrast in the computed tomography (CT) volume data. This paper proposes a consistent semi-automatic landmarks selection and replacing procedure, then uses thin-plate splines to deform a 3D geometric prior model to match the 3D patient CT volume, producing a "best-fit" patient specific polygonal mesh of the whole tooth.
Ruofeng Tong 0001, Jinxiang Dong
CSCWD (1)2
2005 A parallel algorithm of polygons packing based on ant colony
abstract
This paper presents a novel algorithm for optimal packing problem by combining ant colony algorithm with BLF (bottom-left-fill) heuristic approach. The proposed algorithm not only automatically looks for the best sequence of the polygons and each polygon's optimum rotation by ant colony algorithm but also implements the exact layout with the BLF heuristic algorithm. Moreover, the algorithm supports the parallel computation and facilitates quick convergence to the optimal solution. The experimental results show the effectiveness of our algorithm comparing with the other methods.
Wen Peng, Ruofeng Tong 0001, Min Tang 0001, Jinxiang Dong
CSCWD (2)2
2005 Octree-based camera planning in virtual environments
abstract
When people navigate through virtual environments, they often manually control the camera, which may result in dizzy camera motions. In this paper, we introduce a new algorithm for camera planning in virtual environments, which will complete path planning, trajectory planning and orientation planning of virtual camera for the user with desired aesthetic qualities, while the user simply specifies the goal position. The path planning part of the algorithm is based on octree representation of the scene, associating with some ideas from visibility graph. It is more efficient and the resulting path is suitable for camera motion. We also describe an innovative approach for path smoothing and camera speed control in sharp turns, which strengthens the aesthetic effect with little overhead.
Yi-bin Zhang, Ruofeng Tong 0001, Jinxiang Dong
CSCWD (2)2
2005 Brep model simplification for feature suppressing using local error evaluation
abstract
The CAD model simplification technology has been paid more and more attentions for the requirement of seamless integration of CAD/CAE/CAM. This paper presents a new approach for feature suppressing of Brep models using local error evaluation. The improvement is that several features can synchronously be suppressed, and time overhead will be reduced at initial step without recognizing among most of them.
Ruofeng Tong 0001, Tianyang Dong, Jinxiang Dong
CSCWD (2)2
2005 A collaborative approach to assembly sequence planning
Tianyang Dong, Ruofeng Tong 0001, Jinxiang Dong
Adv. Eng. Informatics2
2002 A Model and Algorithm of Two-dimensional Optimum Layout in Blanking
abstract
In this paper, a model and algorithm of optimum layout is proposed for blank layout of single pattern on single rectangular sheet. The approach arranges the cutting patterns on rectangular sheet in the typical style of Double Opposite Layout (DOL). The remains of the sheet, which has been left in one direction after arranging patterns, called as "Step Leavings" (SL) is considered while the optimum layout is discussed We make full use of material of the sheet, including the A in x-direction and y-direction, and set up a reasonable mathematical model for the two-dimensional optimum layout. A corresponding algorithm is also provided to work out the optimum layout quickly and automatically. In this way, we can get the best layout scheme to arrange the maximum shapes on the sheet.
Min Tang 0001, Ruofeng Tong 0001, Jinxiang Dong
CSCWD3
2002 A Feature-Based Collaborative CAD System
abstract
With the intensification of the competition in manufacture, the distributed technology, whose aim is to promote product design process, has changed the traditional CAD serial design approach. But the distributed design systems also bring some new problems such as design conflict. To avoid this inconsistent situation, there must be some coordination mechanisms. At the same time, these mechanisms must not constrain the freedom of the designers too much to take their creativity away. This paper introduces a feature-based distributed CAD system to support this collaborative work. We analyze the reason of the design conflict. For addressing this conflict, we present feature-based concurrency operation model. In such model, the feature is the basic atom that can be locked and excluded from other designers using. Comparing with part level concurrency system, this mechanism doesn't limit the design flexibility too much.
Liangjun Zhang, Min Tang 0001, Ruofeng Tong 0001, Jinxiang Dong
CSCWD3
2002 A Mesh Watermarking Approach for Appearance Attributes
abstract
We describe an algorithm to watermark appearance attributes, as well as the shape of the mesh. Appearance attributes are potential watermarking primitives and the watermarking approach for them can be generalized from that for the shape. The major challenge of generalization is that the watermarking for appearance attributes has more constraints. We focus on this challenge. Especially for the normal vector, we embed the watermark by modifying its orientation, not magnitude. Results show our scheme effectively improves the capacity and enhances the robustness of mesh watermarking.
Liangjun Zhang, Ruofeng Tong 0001, Feiqi Su, Jinxiang Dong
PG2
2002 A Hybrid Model for Smoke Simulation
Ruofeng Tong 0001, Jinxiang Dong
J. Comput. Sci. Technol.1
2002 A volume-preserving approach for modeling and animating water flows generated by metaballs
Ruofeng Tong 0001, Kazufumi Kaneda, Hideo Yamashita
Vis. Comput.1