Jiamin Xu

dblp:155/3294 · DBLP profile ↗
← Back
34ranked-venue papers
16as first author
29since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 7 first-author · 12 since 2021Artificial intelligence and machine learning · 10 · 8 first-author · 7 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Labeling-free RAG-enhanced LLM for intelligent fault diagnosis via reinforcement learning
Jiamin Xu, Zhaohui Jiang 0001, Zhiwen Chen 0001, Hao Luo 0003, Yalin Wang 0003, Weihua Gui 0001
Adv. Eng. Informatics1
2026 Quantum Support Vector Machines and Quantum Kernel Methods
abstract
Background Quantum support vector machines and quantum kernel methods have emerged as promising approaches within quantum machine learning, with the goal of leveraging quantum computing to enhance classification performance and computational efficiency. This review systematically surveys recent advances in QSVM and kernel‐based quantum classifiers, and analyzes their algorithmic frameworks, experimental implementations, and practical challenges. Methods We systematically examine QSVM approaches, including three schemes based on HHL algorithm, Hadamard test, and variational optimization, alongside multi‐class extension and quantum feature mapping. Result Findings indicate that QSVM can offer theoretical speed‐ups in specific settings, particularly when combined with quantum feature mappings that encode data into high‐dimensional Hilbert spaces. However, current implementations remain constrained by hardware limitations, and a lack of large‐scale general validation. Most studies focus on proof‐of‐principle experiments with limited real‐world applicability. Conclusion While promising, QSVM requires further work on scalability, noise resilience, and real‐world integration. Future research should focus on robust algorithms and empirical studies.
Jiamin Xu, Weishan Zhang
Softw. Pract. Exp.4
2026 A 1024-Ch 583-nW/Ch Spike-Sorting SoC With Sparsity-Aware Spike Detection Scratchpad and Ultra-Low-Leakage Dual-Voltage 5T-SRAM for 16K-Template Clustering
abstract
This paper presents an energy-efficient spike-sorting system-on-chip (SoC) designed for closed-loop brain-computer interfaces of massive probing channels. The design first incorporates a sparsity/similarity-aware spike detection scratchpad, leveraging a bit-wise differential encoder and zero-friendly read-out circuits, reducing the dynamic power consumption of spike detection by 77.7%. To mitigate static power dissipation, it also introduces an ultra-low-leakage dual-voltage 5T-SRAM array with level-shifter embedded sense amplifiers, achieving an 82.2% leakage power reduction of neural signal buffering by applying half$V_{DD}$on SRAM cells. Additionally, a memory hierarchy architecture combining on-chip SRAM and off-chip FeRAM, along with a firing-rate-based Osort for cluster template management, minimizes off-chip memory access to only 9.7% with a latency of$11.7\mu $s for 1024-channel spike sorting. A silicon prototype is fabricated in 28-nm CMOS technology, which achieves a power consumption of 583nW/channel and an area consumption of 0.0012mm2/channel. The chip supports real-time spike sorting with up to 16K templates,$21.3\times $greater than the state-of-the-art spike-sorting processor.
Hao Jiang 0024, Zexing Chen, Jiajun Lu, Siqi He, Liangjian Lyu, Jiamin Xu, Shiwei Liu 0002, Yingping Chen, Chixiao Chen, Qi Liu 0010, Ming Liu 0022
IEEE Trans. Circuits Syst. I Regul. Pap.6
2025 OmniSR: Shadow Removal Under Direct and Indirect Lighting
abstract
Shadows can originate from occlusions in both direct and indirect illumination. Although most current shadow removal research focuses on shadows caused by direct illumination, shadows from indirect illumination are often just as pervasive, particularly in indoor scenes. A significant challenge in removing shadows from indirect illumination is obtaining shadow-free images to train the shadow removal network. To overcome this challenge, we propose a novel rendering pipeline for generating shadowed and shadow-free images under direct and indirect illumination, and create a comprehensive synthetic dataset that contains over 30,000 image pairs, covering various object types and lighting conditions. We also propose an innovative shadow removal network that explicitly integrates semantic and geometric priors through concatenation and attention mechanisms. The experiments show that our method outperforms state-of-the-art shadow removal techniques and can effectively generalize to indoor and outdoor scenes under various lighting conditions, enhancing the overall effectiveness and applicability of shadow removal methods.
Jiamin Xu, Renshu Gu, Weiwei Xu 0003, Gang Xu 0001
AAAI1
2025 Detail-Preserving Latent Diffusion for Stable Shadow Removal
abstract
Achieving high-quality shadow removal with strong generalizability is challenging in scenes with complex global illumination. Due to the limited diversity in shadow removal datasets, current methods are prone to overfitting training data, often leading to reduced performance on unseen cases. To address this, we leverage the rich visual priors of a pre-trained Stable Diffusion (SD) model and propose a two-stage fine-tuning pipeline to adapt the SD model for stable and efficient shadow removal. In the first stage, we fix the VAE and fine-tune the denoiser in latent space, which yields substantial shadow removal but may lose some high-frequency details. To resolve this, we introduce a second stage, called the detail injection stage. This stage selectively extracts features from the VAE encoder to modulate the decoder, injecting fine details into the final results. Experimental results show that our method outperforms state-of-the-art shadow removal techniques. The cross-dataset evaluation further demonstrates that our method generalizes effectively to unseen data, enhancing the applicability of shadow removal methods.
Jiamin Xu, Chi Wang 0004, Renshu Gu, Weiwei Xu 0003, Gang Xu 0001
CVPR1
2025 Multisource Knowledge Retrieval Augmented LLM-based Fault Diagnosis Method for Traction Drive Systems
abstract
Fault diagnosis is crucial for ensuring the safety and reliable operation of train traction drive systems. Traditional methods primarily depend on expert experience and historical fault records for analysis. However, as knowledge sources and types expand, existing approaches struggle to integrate multi-source heterogeneous knowledge effectively, which compromises the accuracy of fault diagnosis and the formulation of maintenance decisions. To address these challenges, this paper proposes a multisource knowledge retrieval-augmented large language model (LLM)-based fault diagnosis method. First, a multisource retrieval mechanism is introduced, which uses vector databases and knowledge graphs to extract highly relevant unstructured text and structured knowledge. Next, a reranking model refines the retrieved information, which ensures that the large language model accesses high-quality reference knowledge. Finally, prompt templates integrate optimized textual and structured knowledge that guide the LLM to generate accurate and interpretable fault diagnosis responses. Experimental results show that the proposed method significantly improves the accuracy and reliability of fault diagnosis for traction drive systems while improving the applicability of large language models.
Zhiwen Chen 0001, Jiamin Xu, Lingli Tan, Yuri A. W. Shardt
IECON3
2025 Efficient Object Reconstruction with Differentiable Area Light Shading
abstract
In 3D object reconstruction from photographs, estimating material properties is challenging. We propose an inverse rendering method that uses active area lighting: as this provides a wider range of lighting angles per photo than point lighting, material reconstruction can be more accurate for the same number of photos. We compare area light shading with point lighting. With either mesh or 3D Gaussian splatting pipelines, area lighting can improve BRDF reconstruction and leads to +3 dB relighting PSNR over point lights, or need only \(\nicefrac {1}{5}\) of the input photos for the same quality. We also compare area light shading with Monte Carlo ray tracing and with differential linearly transformed cosines (LTC) plus shadow visibility weighting. LTC can be faster, improving optimization times by 25%. In SOTA method-level comparisons, our approach improves material reconstruction, particularly for material roughness, leading to superior relighting quality.
Yaoan Gao, Jiamin Xu, James Tompkin 0001, Qi Wang 0111, Hujun Bao, Yujun Shen, Huamin Wang 0001, Changqing Zou, Weiwei Xu 0003
SIGGRAPH Asia2
2025 Hybrid Mesh-Neural Representation for 3D Transparent Object Reconstruction
abstract
In this study, we propose a novel method to reconstruct the 3D shapes of transparent objects using images captured by handheld cameras under natural lighting conditions. It combines the advantages of an explicit mesh and multi-layer perceptron (MLP) network as a hybrid representation to simplify the capture settings used in recent studies. After obtaining an initial shape through multi-view silhouettes, we introduced surface-based local MLPs to encode the vertex displacement field (VDF) for reconstructing surface details. The design of local MLPs allowed representation of the VDF in a piecewise manner using two-layer MLP networks to support the optimization algorithm. Defining local MLPs on the surface instead of on the volume also reduced the search space. Such a hybrid representation enabled us to relax the ray-pixel correspondences that represent the light path constraint to our designed ray-cell correspondences, which significantly simplified the implementation of a single-image-based environment-matting algorithm. We evaluated our representation and reconstruction algorithm on several transparent objects based on ground truth models. The experimental results show that our method produces high-quality reconstructions that are superior to those of state-of-the-art methods using a simplified data-acquisition setup.
Jiamin Xu, Zihan Zhu, Hujun Bao, Weiwei Xu 0003
Comput. Vis. Media1
2025 A novel two-stage variables contribution analysis method toward explainable graph convolutional network-based industrial fault diagnosis
Jiamin Xu, Siwen Mo, Zhiwen Chen 0001, Haobin Ke, Zhaohui Jiang 0001
Eng. Appl. Artif. Intell.1
2025 Interpretable procedural material graph generation via diffusion models from reference images
Xiaoyu Lv, Zizhao Wu, Jiamin Xu, Xiaoling Gu, Ming Zeng 0008, Weiwei Xu 0003
Vis. Comput.3
2024 Tabletop Scene Synthesis via Diffusion Models
abstract
Tabletop scene synthesis is an important but underexplored area in scene generation. Existing methods typically focus on larger furniture pieces, often neglecting the fine-grained details of smaller objects positioned on surfaces like tabletops. To address this gap, we propose the first method specifically designed for tabletop scene synthesis using a diffusion model. Our diffusion model, trained on a large dataset of tabletop scenes, achieves state-of-the-art performance compared to current scene generation methods.
Qihuang Zheng, Jiamin Xu
CW2
2024 Gypsophila: A Scalable and Bandwidth-Optimized Multi-Scalar Multiplication Architecture
abstract
Multi-Scalar Multiplication (MSM) is a fundamental cryptographic primitive, which plays a crucial role in Zero-knowledge proof systems. In this paper, we optimize the single MSM Process Element (PE) utilizing buckets with fewer conflicts, enhanced by Greedy-based scheduling, to achieve higher efficiency. The evaluation results show our optimized single MSM PE achieving a speedup of over two times on average, peaking at 3.63 times compared to previous works. Furthermore, we introduce Gypsophila, a scalable and bandwidth-optimized architecture for implementing multiple MSM PEs. Leveraging the characteristics of the bucket method, we optimize the data flow by balancing the throughput of bucket classification, bucket aggregation, and result aggregation in MSM. Simultaneously, multiple PEs with different data access patterns share a universal point input channel and post-processing unit, which improves the module utilization and mitigates the bandwidth pressure. Gypsophila with 16 PEs, accomplishes 16 MSM tasks in a mere 1.01% additional time, showcasing an approximate 7.8% reduction in area, with only about 116 of the bandwidth requirement, compared with 16 PEs without input channel and post-process unit sharing.
Changxu Liu, Hao Zhou 0015, Jiamin Xu, Patrick Dai, Fan Yang 0001
DAC4
2024 3D Human Pose Estimation from Multiple Dynamic Views via Single-view Pretraining with Procrustes Alignment
abstract
3D Human pose estimation from multiple cameras with unknown calibration has received less attention than it should. The few existing data-driven solutions do not fully exploit 3D training data that are available on the market, and typically train from scratch for every novel multi-view scene, which impedes both accuracy and efficiency. We show how to exploit 3D training data to the fullest and associate multiple dynamic views efficiently to achieve high precision on novel scenes using a simple yet effective framework, dubbed Multiple Dynamic View Pose estimation (MDVPose). MDVPose utilizes novel scenarios data to finetune a single-view pretrained motion encoder in multi-view setting, aligns arbitrary number of views in a unified coordinate via Procruste alignment, and imposes multi-view consistency. The proposed method achieves 22.1 mm P-MPJPE or 34.2 mm MPJPE on the challenging in-the-wild Ski-Pose PTZ dataset, which outperforms the state-of-the-art method by 24.8% P-MPJPE (-7.3 mm) and 19.0% MPJPE (-8.0 mm). It also outperforms the state-of-the-art methods by a large margin (-18.2mm P-MPJPE and -28.3mm MPJPE) on the EgoBody dataset. In addition, MDVPose achieves robust performance on the Human3.6M datasets featuring multiple static cameras. Code is available at https://github.com/iGame-Lab/MDVPose.
Renshu Gu, Yixuan Si, Fei Gao 0006, Jiamin Xu, Gang Xu 0001
ACM Multimedia5
2024 Local Gaussian Density Mixtures for Unstructured Lumigraph Rendering
abstract
PSNR 29.47 PSNR 28.80 PSNR 27.28 PSNR 30.84 PSNR 26.69 PSNR 26.
Xiuchao Wu, Jiamin Xu, Chi Wang 0004, Yifan Peng 0001, Qixing Huang, James Tompkin 0001, Weiwei Xu 0003
SIGGRAPH Asia2
2024 OHCA-GCN: A novel graph convolutional network-based fault diagnosis method for complex systems via supervised graph construction and optimization
Jiamin Xu, Haobin Ke, Zhaohui Jiang 0001, Siwen Mo, Zhiwen Chen 0001, Weihua Gui 0001
Adv. Eng. Informatics1
2024 A novel positive-negative graph convolutional network-based fault diagnosis method with application to complex systems
Jiamin Xu, Siwen Mo, Zhaohui Jiang 0001, Zhiwen Chen 0001, Weihua Gui 0001
Neurocomputing1
2024 Blind quality evaluator for multi-exposure fusion image via joint sparse features and complex-wavelet statistical characteristics
Benquan Yang, Yueli Cui, Lihong Liu, Jiamin Xu
Multim. Syst.5
2024 SPH-Net: Hyperspectral Image Super-Resolution via Smoothed Particle Hydrodynamics Modeling
abstract
Reconstructing a high-resolution hyperspectral image (HSI) from a low-resolution HSI is significant for many applications, such as remote sensing and aerospace. Most deep learning-based HSI super-resolution methods pay more attention to developing novel network structures but rarely study the HSI super-resolution problem from the perspective of image dynamic evolution. In this article, we propose that the HSI pixel motion during the super-resolution reconstruction process can be analogized to the particle movement in the smoothed particle hydrodynamics (SPH) field. To this end, we design an SPH network (SPH-Net) for HSI super-resolution in light of the SPH theory. Specifically, we construct a smooth function based on SPH and design a smooth convolution in multiscales to exploit spectral correlation and preserve the spectral information in the super-resolved image. In addition, we apply the SPH approximation method to discretize the Navier-Stokes motion equation into SPH equation form, which can guide the HSI pixel motion in the desired direction during super-resolution reconstruction, thereby producing clear edges in the spatial domain. Experiments on three public hyperspectral datasets demonstrate that the proposed SPH-Net outperforms the state-of-the-art methods in terms of objective metrics and visual quality.
Mingjin Zhang, Jiamin Xu, Jing Zhang 0037, Haimei Zhao, Wenteng Shang, Xinbo Gao 0001
IEEE Trans. Cybern.2
2023 Hybrid Position Sensorless Control Based on Estimation Position Error Switching for PMSM in Full Speed Range
abstract
In permanent magnet synchronous motor (PMSM) full speed range position sensorless control, the conventional hybrid control method usually determines the switching speed point as 10%-30% of the rated speed in the switching process empirically, which fails to effectively reduce the pulsations of position and speed. In this paper, a method is proposed to determine the switching speed point based on the scalarized estimated position error to achieve smooth switching and reduce the pulsations of position and speed during switching. The method is divided into three ranges: In zero-low and medium-high speed ranges, the estimated speed and position are obtained using the high-frequency injection method and the sliding mode observer approach, respectively. In transition range, the switching speed is determined when the position errors of the two methods are similar. After weighing the two errors, the estimated position and speed are acquired through a phase-locked loop. Finally, the effectiveness of the proposed method is verified in MATLAB/Simulink based on the built-in PMSM.
Xinran Shi, Jinglin Liu, Jiamin Xu
IECON3
2023 Proportional Resonant Filtering For Improved SMO With Optimized Critical Saturation Switching Function
abstract
This paper proposes a sensorless speed control strategy for a permanent-magnet synchronous motor (PMSM) based on proportional resonant filtering and an improved sliding-mode observer (SMO) with an optimized critical saturation switching function. This strategy suppresses pulsation and decreases the delay to improve the accuracy of position estimation. First, an optimized critical saturation switching function was designed to output a sine wave with a boundary layer fixed at 1. This strategy kept the convergence speed of SMO constant and weakened the pulsation. However, the phase delay in the estimated back-EMF caused by introducing a low-pass filter (LPF) persists in the system. Therefore, this paper proposes a proportional resonant filter (PRF) to address this problem. The PRF, without phase delay, can extract the fundamental component and filter out the harmonic components in the estimated back electromotive force (back-EMF). The simulation results verify the effectiveness and accuracy of the proposed strategy.
Jiamin Xu, Jinglin Liu, Xinran Shi
IECON1
2023 Robust twin depth support vector machine based on average depth
Jiamin Xu, Huamin Wang 0002, Shiping Wen 0001
Knowl. Based Syst.1
2023 Multichannel Domain Adaptation Graph Convolutional Networks-Based Fault Diagnosis Method and With Its Application
abstract
Intelligent fault diagnosis of the complex systems has made great progress based on the availability of massive labeled data. However, due to the diversity of working conditions and the lack of sufficient fault samples in practice, the generalization of the existing fault diagnosis methods are weak. To handle this issue, a multichannel domain adaptation graph convolutional network method is proposed. In the proposed network, a feature mapping layer based on convolutional neural network is used first to extract features from input data, which then are transmitted to the graph generator to construct two association graphs. After that, three distributed graph convolutional networks are used to extract the specific and common embeddings from two association graphs and their combination. Meanwhile, to fuse these embeddings adaptively, an attention mechanism is used to learn importance weights. Besides, a domain discriminator is leveraged to reduce the distribution discrepancy of different data domains. Finally, a label classifier is used to output fault diagnosis results. Two experimental studies with different signal types show that the proposed method not only presents better diagnosis performance than existing methods with few samples, but also can extract domain-invariant features for cross-domain under varying working conditions.
Zhiwen Chen 0001, Haobin Ke, Jiamin Xu, Tao Peng 0010, Chunhua Yang 0001
IEEE Trans. Ind. Informatics3
2023 Oversmoothing Relief Graph Convolutional Network-Based Fault Diagnosis Method With Application to the Rectifier of High-Speed Trains
abstract
In the conventional graph convolutional network (GCN)-based fault diagnosis method, multilayer GCN model is often used for feature extraction. However, the application of multilayer GCN will encounter oversmoothing problem, and thus reduce the diagnostic performance. Therefore, the oversmoothing relief GCN (OsR-GCN) method is proposed. Specifically, two association graph construction methods, namely the Euclidean distance (ED)-based method and the structure analysis (SA)-based method, are first introduced. Then, the constructed graph and measurements are input to the OsR-GCN model, in which a weight coefficient is proposed to relieve the oversmoothing problem. Next, an improved particle swarm optimization algorithm is introduced to find the optimal weight coefficient. Finally, the proposed method is applied to diagnose the pulse rectifier faults in a hardware-in-the-loop simulated traction control system of high-speed trains. The achieved results show that the proposed method outperforms the existing fault diagnosis methods.
Jiamin Xu, Haobin Ke, Zhiwen Chen 0001, Xinyu Fan 0003, Tao Peng 0010, Chunhua Yang 0001
IEEE Trans. Ind. Informatics1
2023 ScaNeRF: Scalable Bundle-Adjusting Neural Radiance Fields for Large-Scale Scene Rendering
abstract
High-quality large-scale scene rendering requires a scalable representation and accurate camera poses. This research combines tile-based hybrid neural fields with parallel distributive optimization to improve bundle-adjusting neural radiance fields. The proposed method scales with a divide-and-conquer strategy. We partition scenes into tiles, each with a multi-resolution hash feature grid and shallow chained diffuse and specular multilayer perceptrons (MLPs). Tiles unify foreground and background via a spatial contraction function that allows both distant objects in outdoor scenes and planar reflections as virtual images outside the tile. Decomposing appearance with the specular MLP allows a specular-aware warping loss to provide a second optimization path for camera poses. We apply the alternating direction method of multipliers (ADMM) to achieve consensus among camera poses while maintaining parallel tile optimization. Experimental results show that our method outperforms state-of-the-art neural scene rendering method quality by 5%--10% in PSNR, maintaining sharp distant objects and view-dependent reflections across six indoor and outdoor scenes.
Xiuchao Wu, Jiamin Xu, Hujun Bao, Qixing Huang, Yujun Shen, James Tompkin 0001, Weiwei Xu 0003
ACM Trans. Graph.2
2022 Graph Convolutional Network-Based Method for Fault Diagnosis Using a Hybrid of Measurement and Prior Knowledge
abstract
Deep-neural network-based fault diagnosis methods have been widely used according to the state of the art. However, a few of them consider the prior knowledge of the system of interest, which is beneficial for fault diagnosis. To this end, a new fault diagnosis method based on the graph convolutional network (GCN) using a hybrid of the available measurement and the prior knowledge is proposed. Specifically, this method first uses the structural analysis (SA) method to prediagnose the fault and then converts the prediagnosis results into the association graph. Then, the graph and measurements are sent into the GCN model, in which a weight coefficient is introduced to adjust the influence of measurements and the prior knowledge. In this method, the graph structure of GCN is used as a joint point to connect SA based on the model and GCN based on data. In order to verify the effectiveness of the proposed method, an experiment is carried out. The results show that the proposed method, which combines the advantages of both SA and GCN, has better diagnosis results than the existing methods based on common evaluation indicators.
Zhiwen Chen 0001, Jiamin Xu, Tao Peng 0010, Chunhua Yang 0001
IEEE Trans. Cybern.2
2022 Real-Time Fault Diagnosis of Pulse Rectifier in Traction System Based on Structural Model
abstract
The pulse rectifier of a traction system in locomotives and electric multiple unit (EMUs) is usually vulnerable to performance degradation and faults due to uncertain factors, such as vibrations, aging, electromagnetic interferences. In order to ensure adequate redundancy and isolation measures in time to avoid fault propagation in traction systems, a real-time fault diagnosis method for sensors and IGBTs of the impulse rectifier is proposed, which lays a foundation for the redundant design of traction systems. Once a fault is detected in this paper, which can immediately act to the detected faults and take effective remedial measures to avoid the failure of the whole system. It is based on the structural analysis of the traction system whose structural model of interest will be established. Meanwhile, the structural model is evaluated and optimized according to the analytical relation model under various fault conditions, and the minimum structural overdetermined sets (MSOs) are obtained based on the optimized model, which can be used to isolate all faults. Using the MSOs, the redundancy relationship is deduced and the sequence residuals are generated. Afterward, the cumulative sum (CUSUM) algorithm is used for diagnosis decision making. The effectiveness of the proposed method is finally verified on a hardware-in-loop test platform, which can accurately simulate a traction system. It shows that the proposed method can achieve both good feasibility and high accuracy.
Xueming Li 0003, Jiamin Xu, Zhiwen Chen 0001, Shaolong Xu, Kan Liu 0002
IEEE Trans. Intell. Transp. Syst.2
2022 Scalable neural indoor scene rendering
abstract
We propose a scalable neural scene reconstruction and rendering method to support distributed training and interactive rendering of large indoor scenes. Our representation is based on tiles. Tile appearances are trained in parallel through a background sampling strategy that augments each tile with distant scene information via a proxy global mesh. Each tile has two low-capacity MLPs: one for view-independent appearance (diffuse color and shading) and one for view-dependent appearance (specular highlights, reflections). We leverage the phenomena that complex view-dependent scene reflections can be attributed to virtual lights underneath surfaces at the total ray distance to the source. This lets us handle sparse samplings of the input scene where reflection highlights do not always appear consistently in input images. We show interactive free-viewpoint rendering results from five scenes, one of which covers an area of more than 100 m 2 . Experimental results show that our method produces higher-quality renderings than a single large-capacity MLP and five recent neural proxy-geometry and voxel-based baseline methods. Our code and data are available at project webpage https://xchaowu.github.io/papers/scalable-nisr.
Xiuchao Wu, Jiamin Xu, Zihan Zhu, Hujun Bao, Qixing Huang, James Tompkin 0001, Weiwei Xu 0003
ACM Trans. Graph.2
2021 Sleep Analysis During Light Sleep Based on K-means Clustering and BiLSTM
Jiamin Xu, Haojun Sun
WISA1
2021 Scalable image-based indoor scene rendering with reflections
abstract
This paper proposes a novel scalable image-based rendering (IBR) pipeline for indoor scenes with reflections. We make substantial progress towards three sub-problems in IBR, namely, depth and reflection reconstruction, view selection for temporally coherent view-warping, and smooth rendering refinements. First, we introduce a global-mesh-guided alternating optimization algorithm that robustly extracts a two-layer geometric representation. The front and back layers encode the RGB-D reconstruction and the reflection reconstruction, respectively. This representation minimizes the image composition error under novel views, enabling accurate renderings of reflections. Second, we introduce a novel approach to select adjacent views and compute blending weights for smooth and temporal coherent renderings. The third contribution is a supersampling network with a motion vector rectification module that refines the rendering results to improve the final output's temporal coherence. These three contributions together lead to a novel system that produces highly realistic rendering results with various reflections. The rendering quality outperforms state-of-the-art IBR or neural rendering algorithms considerably.
Jiamin Xu, Xiuchao Wu, Zihan Zhu, Qixing Huang, Yin Yang 0002, Hujun Bao, Weiwei Xu 0003
ACM Trans. Graph.1
2019 Survey of 3D modeling using depth cameras
abstract
Three-dimensional (3D) modeling is an important topic in computer graphics and computer vision. In recent years, the introduction of consumer-grade depth cameras has resulted in profound advances in 3D modeling. Starting with the basic data structure, this survey reviews the latest developments of 3D modeling based on depth cameras, including research works on camera tracking, 3D object and scene reconstruction, and high-quality texture reconstruction. We also discuss the future work and possible solutions for 3D modeling based on the depth camera.
Hantong Xu, Jiamin Xu, Weiwei Xu 0003
Virtual Real. Intell. Hardw.2
2018 Online Global Non-rigid Registration for 3D Object Reconstruction Using Consumer-level Depth Cameras
abstract
Abstract We investigate how to obtain high‐quality 360‐degree 3D reconstructions of small objects using consumer‐level depth cameras. For many homeware objects such as shoes and toys with dimensions around 0.06 – 0.4 meters, their whole projections, in the hand‐held scanning process, occupy fewer than 20% pixels of the camera's image. We observe that existing 3D reconstruction algorithms like KinectFusion and other similar methods often fail in such cases even under the close‐range depth setting. To achieve high‐quality 3D object reconstruction results at this scale, our algorithm relies on an online global non‐rigid registration, where embedded deformation graph is employed to handle the drifting of camera tracking and the possible nonlinear distortion in the captured depth data. We perform an automatic target object extraction from RGBD frames to remove the unrelated depth data so that the registration algorithm can focus on minimizing the geometric and photogrammetric distances of the RGBD data of target objects. Our algorithm is implemented using CUDA for a fast non‐rigid registration. The experimental results show that the proposed method can reconstruct high‐quality 3D shapes of various small objects with textures.
Jiamin Xu, Weiwei Xu 0003, Yin Yang 0002, Zhigang Deng 0001, Hujun Bao
Comput. Graph. Forum1
2016 A new method for multi-oriented graphics-scene-3D text classification in video
Jiamin Xu, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan, Seiichi Uchida
Pattern Recognit.1
2014 Graphics and Scene Text Classification in Video
abstract
Achieving good accuracy for text detection and recognition is a challenging and interesting problem in the field of video document analysis because of the presences of both graphics text that has good clarity and scene text that is unpredictable in video frames. Therefore, in this paper, we present a novel method for classifying graphics texts and scene texts by exploiting temporal information and finding the relationship between them in video. The method proposes an iterative procedure to identify Probable Graphics Text Candidates (PGTC) and Probable Scene Text Candidates (PSTC) in video based on the fact that graphics texts in general do not have large movements especially compared to scene texts which are usually embedded on background. In addition to PGTC and PSTC, the iterative process automatically identifies the number of video frames with the help of a converging criterion. The method further explores the symmetry between intra and inter character components to identify graphics text candidates and scene text candidates. Boundary growing method is employed to restore the complete text line. For each segmented text line, we finally introduce Eigen value analysis to classify graphics and scene text lines based on the distribution of respective Eigen values. Experimental results with the existing methods show that the proposed method is effective and useful to improve the accuracy of text detection and recognition.
Jiamin Xu, Palaiahnakote Shivakumara, Tong Lu 0002, Trung Quy Phan, Chew Lim Tan
ICPR1
2014 2D and 3D Video Scene Text Classification
abstract
Text detection and recognition is a challenging problem in document analysis due to the presence of the unpredictable nature of video texts, such as the variations of orientation, font and size, illumination effects, and even different 2D/3D text shadows. In this paper, we propose a novel horizontal and vertical symmetry feature by calculating the gradient direction and the gradient magnitude of each text candidate, which results in Potential Text Candidates (PTCs) after applying the k-means clustering algorithm on the gradient image of each input frame. To verify PTCs, we explore temporal information of video by proposing an iterative process that continuously verifies the PTCs of the first frame and the successive frames, until the process meets the converging criterion. This outputs Stable Potential Text Candidates (SPTCs). For each SPTC, the method obtains text representatives with the help of the edge image of the input frame. Then for each text representative, we divide it into four quadrants and check a new Mutual Nearest Neighbor Symmetry (MNNS) based on the dominant stroke width distances of the four quadrants. A voting method is finally proposed to classify each text block as either 2D or 3D by counting the text representatives that satisfy MNNS. Experimental results on classifying 2D and 3D text images are promising, and the results are further validated by text detection and recognition before classification and after classification with the exiting methods, respectively.
Jiamin Xu, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan
ICPR1