Jianwen Lou

dblp:189/4522 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 3DTeethSAM: Taming SAM2 for 3D Teeth Segmentation
abstract
3D teeth segmentation, involving the localization of tooth instances and their semantic categorization in 3D dental models, is a critical yet challenging task in digital dentistry due to the complexity of real-world dentition. In this paper, we propose 3DTeethSAM, an adaptation of the Segment Anything Model 2 (SAM2) for 3D teeth segmentation. SAM2 is a pretrained foundation model for image and video segmentation, demonstrating a strong backbone in various downstream scenarios. To adapt SAM2 for 3D teeth data, we render images of 3D teeth models from predefined views, apply SAM2 for 2D segmentation, and reconstruct 3D results using 2D-3D projections. Since SAM2's performance depends on input prompts and its initial outputs often have deficiencies, and given its class-agnostic nature, we introduce three light-weight learnable modules: (1) a prompt embedding generator to derive prompt embeddings from image embeddings for accurate mask decoding, (2) a mask refiner to enhance SAM2's initial segmentation results, and (3) a mask classifier to categorize the generated masks. Additionally, we incorporate Deformable Global Attention Plugins (DGAP) into SAM2's image encoder. The DGAP enhances both the segmentation accuracy and the speed of the training process. Our method has been validated on the 3DTeethSeg benchmark, achieving an IoU of 91.90% on high-resolution 3D teeth meshes, establishing a new state-of-the-art in the field.
Zhiguo Lu, Jianwen Lou, Hairong Jin, Youyi Zheng, Kun Zhou 0001
AAAI2
2025 Learning center- and boundary-aware instance representation for 3D tooth segmentation
Hairong Jin, Jianwen Lou, Zhiguo Lu, Kun Zhou 0001, Youyi Zheng
Comput. Graph.2
2025 TSRNet: A Dual-Stream Network for Refining 3D Tooth Segmentation
abstract
The field of 3D tooth segmentation has made considerable advances thanks to deep learning, but challenges remain with coarse segmentation boundaries and prediction errors. In this article, we introduce a novel learnable method to refine coarse results obtained from existing 3D tooth segmentation algorithms. The refinement framework features a dual-stream network called TSRNet (Tooth Segmentation Refinement Network) to rectify defective boundary and distance maps extracted from the coarse segmentation. The boundary map provides explicit boundary information, while the distance map provides gradient information in the form of the shortest geodesic distance between the vertex and the segmentation boundary. Following well-designed rules, the two refined maps are utilized to move the coarse tooth boundaries toward their correct positions through an iterative refinement process. The two-stage refinement method is validated on both 3D tooth and segmentation benchmark datasets. Extensive experiments demonstrate that our method significantly improves upon the coarse results from baseline methods and achieves state-of-the-art performance.
Hairong Jin, Yuefan Shen, Jianwen Lou, Kun Zhou 0001, Youyi Zheng
IEEE Trans. Vis. Comput. Graph.3
2025 Neural Orthodontic Staging: Predicting Teeth Movements With a Transformer
abstract
We present a novel learning-based method for predicting tooth movements in orthodontic treatment path planning (orthodontic staging). Recognizing the multi-solution nature of orthodontic staging, our approach involves generating the staging sequence progressively with a dedicated Transformer model. This model predicts teeth movements within a predefined number of steps (e.g., 10 or 20), targeting alignment in problematic dentition. The Transformer refines its predictions iteratively, building on previous outcomes until reaching a state that aligns with the target within an acceptable distance. This mirrors real-life scenarios where orthodontists dynamically adjust staging plans based on treatment outcomes. Our Transformer model is tailored to incorporate spatial and temporal attentions, addressing inter-tooth and inter-step interactions, respectively. These attentions are further refined with relative positional encoding. Recognizing the significant influence of tooth shape on the alignment process, we propose integrating a tooth-wise shape encoder to extract morphological features from the 3D teeth point cloud. These features are then fused into the Transformer, facilitating the capture of inter-tooth dynamics during staging, in collaboration with spatial attention. We validate the proposed method on a large-scale dataset that contains 10K real-life orthodontic cases. The results show that our method outperforms the state-of-the-art, and orthodontists favor its predictions.
Jiayue Ma, Jianwen Lou, Borong Jiang, Hengyi Ye, Wenke Yu, Xiang Chen 0001, Kun Zhou 0001, Youyi Zheng
IEEE Trans. Vis. Comput. Graph.2
2024 KeypointDETR: An End-to-End 3D Keypoint Detector
Hairong Jin, Yuefan Shen, Jianwen Lou, Kun Zhou 0001, Youyi Zheng
ECCV (74)3
2024 Automatic Indoor Lighting Generation Driven by Human Activity Learned from Virtual Experience
abstract
A good indoor lighting solution should fit with people’s habitual activity and have a low energy cost. However, it’s challenging to capture and model human activity in reality due to its high complexity, let alone incorporating it into lighting planning. As a result, indoor lighting designing still relies on professional’s hands, which is laborious and inefficient. To solve this problem, we propose a novel framework for automatic indoor lighting generation driven by human activity learned from virtual experience. We first harnesses Virtual Reality to simulate and model the user’s daily activities within an indoor scene, and then devises a robust objective function which encompasses multiple activity-driven cost terms for lighting layout optimization. With the objective function and the collected user behavioral data, such as trajectory and head pose, an optimization algorithm is applied to search for the optimal solution. Experiments under different indoor scenes demonstrate that the proposed method can generate lighting solutions that meet personalized behavioral needs in an energy-economic way, which are competitive against those designed by professionals.
Jianwen Lou, Youyi Zheng, Kun Zhou 0001
VR2
2024 Single depth image 3D face reconstruction via domain adaptive learning
Xiaoxu Cai, Jianwen Lou, Jiajun Bu, Junyu Dong, Haishuai Wang, Hui Yu 0001
Frontiers Comput. Sci.2
2024 Perceptual loss guided Generative adversarial network for saliency detection
Xiaoxu Cai, Gaige Wang, Jianwen Lou, Muwei Jian, Junyu Dong, Rung Ching Chen, Brett Stevens, Hui Yu 0001
Inf. Sci.3
2021 Real-Time 3D Facial Tracking via Cascaded Compositional Learning
abstract
We propose to learn a cascade of globally-optimized modular boosted ferns (GoMBF) to solve multi-modal facial motion regression for real-time 3D facial tracking from a monocular RGB camera. GoMBF is a deep composition of multiple regression models with each is a boosted ferns initially trained to predict partial motion parameters of the same modality, and then concatenated together via a global optimization step to form a singular strong boosted ferns that can effectively handle the whole regression target. It can explicitly cope with the modality variety in output variables, while manifesting increased fitting power and a faster learning speed comparing against the conventional boosted ferns. By further cascading a sequence of GoMBFs (GoMBF-Cascade) to regress facial motion parameters, we achieve competitive tracking performance on a variety of in-the-wild videos comparing to the state-of-the-art methods which either have higher computational complexity or require much more training data. It provides a robust and highly elegant solution to real-time 3D facial tracking using a small set of training data and hence makes it more practical in real-world applications. We further deeply investigate the effect of synthesized facial images on training non-deep learning methods such as GoMBF-Cascade for 3D facial tracking. We apply three types synthetic images with various naturalness levels for training two different tracking methods, and compare the performance of the tracking models trained on real data, on synthetic data and on a mixture of data. The experimental results indicate that, i) the model trained purely on synthetic facial imageries can hardly generalize well to unconstrained real-world data, ii) involving synthetic faces into training benefits tracking in some certain scenarios but degrades the tracking model's generalization ability. These two insights could benefit a range of non-deep learning facial image analysis tasks where the labelled real data is difficult to acquire.
Jianwen Lou, Xiaoxu Cai, Junyu Dong, Hui Yu 0001
IEEE Trans. Image Process.1
2020 Hybrid regression and isophote curvature for accurate eye center localization
abstract
Abstract The eye center localization is a crucial requirement for various human-computer interaction applications such as eye gaze estimation and eye tracking. However, although significant progress has been made in the field of eye center localization in recent years, it is still very challenging for tasks under the significant variability situations caused by different illumination, shape, color and viewing angles. In this paper, we propose a hybrid regression and isophote curvature for accurate eye center localization under low resolution. The proposed method first applies the regression method, which is called Supervised Descent Method (SDM), to obtain the rough location of eye region and eye centers. SDM is robust against the appearance variations in the eye region. To make the center points more accurate, isophote curvature method is employed on the obtained eye region to obtain several candidate points of eye center. Finally, the proposed method selects several estimated eye center locations from the isophote curvature method and SDM as our candidates and a SDM-based means of gradient method further refine the candidate points. Therefore, we combine regression and isophote curvature method to achieve robustness and accuracy. In the experiment, we have extensively evaluated the proposed method on the two public databases which are very challenging and realistic for eye center localization and compared our method with existing state-of-the-art methods. The results of the experiment confirm that the proposed method outperforms the state-of-the-art methods with a significant improvement in accuracy and robustness and has less computational complexity.
Jianwen Lou, Junyu Dong, Lin Qi 0004, Gongfa Li, Hui Yu 0001
Multim. Tools Appl.2
2020 Realistic Facial Expression Reconstruction for VR HMD Users
abstract
We present a system for sensing and reconstructing facial expressions of the virtual reality (VR) head-mounted display (HMD) user. The HMD occludes a large portion of the user's face, which makes most existing facial performance capturing techniques intractable. To tackle this problem, a novel hardware solution with electromyography (EMG) sensors being attached to the headset frame is applied to track facial muscle movements. For realistic facial expression recovery, we first reconstruct the user's 3D face from a single image and generate the personalized blendshapes associated with seven facial action units (AUs) on the most emotionally salient facial parts (ESFPs). We then utilize pre-processed EMG signals for measuring activations of AU-coded facial expressions to drive pre-built personalized blendshapes. Since facial expressions appear as important nonverbal cues of the subject's internal emotional states, we further investigate the relationship between six basic emotions - anger, disgust, fear, happiness, sadness and surprise, and detected AUs using a fern classifier. Experiments show the proposed system can accurately sense and reconstruct high-fidelity common facial expressions while providing useful information regarding the emotional state of the HMD user.
Jianwen Lou, Yiming Wang 0001, Charles Nduka, Mahyar Hamedi, Ifigeneia Mavridou, Fei-Yue Wang 0001, Hui Yu 0001
IEEE Trans. Multim.1
2019 Multi-subspace supervised descent method for robust face alignment
abstract
Supervised Descent Method (SDM) is one of the leading cascaded regression approaches for face alignment with state-of-the-art performance and a solid theoretical basis. However, SDM is prone to local optima and likely averages conflicting descent directions. This makes SDM ineffective in covering a complex facial shape space due to large head poses and rich non-rigid face deformations. In this paper, a novel two-step framework called multi-subspace SDM (MS-SDM) is proposed to equip SDM with a stronger capability for dealing with unconstrained faces. The optimization space is first partitioned with regard to shape variations using k-means. The generated subspaces show semantic significance which highly correlates with head poses. Faces among a certain subspace also show compatible shape-appearance relationships. Then, Naive Bayes is applied to conduct robust subspace prediction by concerning about the relative proximity of each subspace to the sample. This guarantees that each sample can be allocated to the most appropriate subspace-specific regressor. The proposed method is validated on benchmark face datasets with a mobile facial tracking implementation.
Jianwen Lou, Xiaoxu Cai, Yiming Wang 0001, Hui Yu 0001, Shaun J. Canavan
Multim. Tools Appl.1
2016 Learning perceptual texture similarity and relative attributes from computational features
abstract
Previous work has shown that perceptual texture similarity and relative attributes cannot be well described by computational features. In this paper, we propose to predict human's visual perception of texture images by learning a non-linear mapping from computational feature space to perceptual space. Hand-crafted features and deep features, which were successfully applied in texture classification tasks, were extracted and used to train Random Forest and rankSVM models against perceptual data from psychophysical experiments. Three texture datasets were used to test our proposed method and the experiments show that the predictions of such learnt models are in high correlation with human's results.
Jianwen Lou, Lin Qi 0004, Junyu Dong, Hui Yu 0001, Guoqiang Zhong 0001
IJCNN1