VLDB 2026 Research / reviewers in the wild / expert
Yihao Luo
dblp:243/6696
· DBLP profile ↗
39ranked-venue papers
10as first author
35since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 5 first-author · 21 since 2021Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explicit differentiable slicing and global deformation for cardiac mesh reconstructionabstractThree-dimensional (3D) mesh reconstruction of the cardiac anatomy from medical images is useful for shape and motion measurements and biophysics simulations. However, 3D medical images are often acquired as 2D slices that are sparsely sampled (e.g., large slice spacing) and noisy, and 3D mesh reconstruction on such data is a challenging task. Traditional voxel-based approaches utilize non-differentiable pre- and post-processing that compromises fidelity to images, while mesh-level deep learning approaches require large 3D mesh annotations that are difficult to obtain. Differentiable cross-domain supervision from 2D images to 3D meshes is therefore crucial for enabling end-to-end optimization in medical imaging. While there have been attempts to approximate the voxelization and slicing of meshes that are being optimized, there has not yet been a method for directly using 2D slices to supervise 3D mesh reconstruction in a differentiable manner. Here, we propose a novel explicit differentiable voxelization and slicing (DVS) algorithm allowing gradient backpropagation to a 3D mesh from its slices, which facilitates refined mesh optimization directly supervised by the losses defined on 2D images. Further, we propose an innovative framework for extracting patient-specific left ventricle (LV) meshes from medical images by coupling DVS with a graph harmonic deformation (GHD) mesh morphing descriptor of cardiac shape that naturally preserves mesh quality and smoothness during optimization. The proposed framework achieves state-of-the-art performance in cardiac mesh reconstruction tasks from densely sampled (CT) as well as sparsely sampled (MRI stack with few slices) images, outperforming alternatives, including Marching Cubes, statistical shape models, algorithms with vertex-based mesh morphing algorithms and alternative methods for image-supervision of mesh reconstruction. Experimental results demonstrate that our method achieves an overall Dice score of 90% during a sparse fitting on multi-datasets. The proposed method can further quantify clinically useful parameters such as ejection fraction and global myocardial strains, closely matching the ground truth and outperforming the traditional voxel-based approach in sparse images. Yihao Luo, Dario Sesia, Fanwen Wang, Yinzhe Wu 0001, Wenhao Ding, Md. Kamrul Hasan 0002, Fadong Shi, Anoop Shah, Amit Kaura, Jamil Mayet, Guang Yang 0006, Choon Hwai Yap |
Medical Image Anal. | 1 |
| 2026 | Multi-view subspace tensorization with attentive clustering embedding
Yanghang Zheng, Haonan Huang, Yihao Luo, Yuning Qiu, Andong Wang, Guoxu Zhou, Qibin Zhao |
Neural Networks | 3 |
| 2026 | 4-D Reconstruction of Fetal Left Ventricle From Echocardiography via 2.5-D Radial Segmentation and Graph-Fourier Reconstruction
Md. Kamrul Hasan 0002, Haziq Shahard, Lucas Iijima, Nida Ruseckaite, Yihao Luo, Iris Scharnreitner, Andreas Tulzer, Bin Liu 0040, Guang Yang 0006, Choon Hwai Yap |
IEEE Trans. Medical Imaging | 6 |
| 2025 | LawDNet: Enhanced Audio-Driven Lip Synthesis via Local Affine Warping DeformationabstractIn the domain of photorealistic talking head generation, the fidelity of audio-driven lip motion synthesis is essential for realistic virtual interactions. Existing methods face two key challenges: a lack of vivacity due to limited diversity in generated lip poses and noticeable anamorphose motions caused by poor temporal coherence. To address these issues, we propose LawD-Net, a novel deep-learning architecture enhancing lip synthesis through a Local Affine Warping Deformation mechanism. This mechanism models the intricate lip movements in response to the audio input by controllable non-linear warping fields. These fields consist of local affine transformations focused on abstract keypoints within deep feature maps, offering a novel universal paradigm for feature warping in networks. Additionally, LawDNet incorporates a dual-stream discriminator for improved frame-to-frame continuity and employs face normalization techniques to handle pose and scene variations. Extensive evaluations demonstrate LawDNet’s superior robustness and lip movement dynamism performance compared to previous methods. Junli Deng, Yihao Luo, Xueting Yang, Siyou Li, Jinyang Guo 0002, Ping Shi 0001 |
ICASSP | 2 |
| 2025 | MeshAnything V2: Artist-Created Mesh Generation with Adjacent Mesh TokenizationabstractMeshes are the de facto 3D representation in the industry but are labor-intensive to produce. Recently, a line of research has focused on autoregressively generating meshes. This approach processes meshes into a sequence composed of vertices and then generates them vertex by vertex, similar to how a language model generates text. These methods have achieved some success but still struggle to generate complex meshes. One primary reason for this limitation is their inefficient tokenization methods. To address this issue, we introduce MeshAnything V2, an advanced mesh generation model designed to create Artist-Created Meshes that align precisely with specified shapes. A key innovation behind MeshAnything V2 is our novel Adjacent Mesh Tokenization (AMT) method. Unlike traditional approaches that represent each face using three vertices, AMT optimizes this by employing a single vertex wherever feasible, effectively reducing the token sequence length by about half on average. This not only streamlines the tokenization process but also results in more compact and well-structured sequences, enhancing the efficiency of mesh generation. With these improvements, MeshAnything V2 effectively doubles the face limit compared to previous models, delivering superior performance without increasing computational costs. We will make our code and models publicly available. Project Page: https://buaacyw.github.io/meshanything-v2/ Yikai Wang 0001, Yihao Luo, Zilong Chen, Jun Zhu 0001, Chi Zhang 0007, Guosheng Lin |
ICCV | 3 |
| 2025 | Two-Stage Generative Model for Intracranial Aneurysm Meshes with Morphological Marker Conditioning
Wenhao Ding, Kangjun Ji, Simão Castro, Yihao Luo, Dylan Roi, Choon Hwai Yap |
MICCAI (10) | 4 |
| 2025 | Sparc3D: Sparse Representation and Construction for High-Resolution 3D Shapes ModelingabstractHigh-fidelity 3D object synthesis remains significantly more challenging than 2D image generation due to the unstructured nature of mesh data and the cubic complexity of dense volumetric grids. Existing two-stage pipelines—compressing meshes with a VAE (using either 2D or 3D supervision), followed by latent diffusion sampling—often suffer from severe detail loss caused by inefficient representations and modality mismatches introduced in VAE.
We introduce Sparc3D, a unified framework that combines a sparse deformable marching cubes representation Sparcubes with a novel encoder Sparconv-VAE. Sparcubes converts raw meshes into high-resolution ($1024^3$) surfaces with arbitrary topology by scattering signed distance and deformation fields onto a sparse cube, allowing differentiable optimization.
Sparconv-VAE is the first modality-consistent variational autoencoder built entirely upon sparse convolutional networks, enabling efficient and near-lossless 3D reconstruction suitable for high-resolution generative modeling through latent diffusion. Sparc3D achieves state-of-the-art reconstruction fidelity on challenging inputs, including open surfaces, disconnected components, and intricate geometry. It preserves fine-grained shape details, reduces training and inference cost, and integrates naturally with latent diffusion models for scalable, high-resolution 3D generation. Yufei Wang 0006, Heliang Zheng, Yihao Luo, Bihan Wen |
NeurIPS | 4 |
| 2025 | Latent low-rank tensor wheel decomposition for visual data completion
Yihao Luo, Yuning Qiu, Hong-Xia Rao, Guoxu Zhou |
Neurocomputing | 1 |
| 2025 | Feedback Attention to Enhance Unsupervised Deep Learning Image Registration in 3D EchocardiographyabstractCardiac motion estimation is important for assessing the contractile health of the heart, and performing this in 3D can provide advantages due to the complex 3D geometry and motions of the heart. Deep learning image registration (DLIR) is a robust way to achieve cardiac motion estimation in echocardiography, providing speed and precision benefits, but DLIR in 3D echo remains challenging. Successful unsupervised 2D DLIR strategies are often not effective in 3D, and there have been few 3D echo DLIR implementations. Here, we propose a new spatial feedback attention (FBA) module to enhance unsupervised 3D DLIR and enable it. The module uses the results of initial registration to generate a co-attention map that describes remaining registration errors spatially and feeds this back to the DLIR to minimize such errors and improve self-supervision. We show that FBA improves a range of promising 3D DLIR designs, including networks with and without transformer enhancements, and that it can be applied to both fetal and adult 3D echo, suggesting that it can be widely and flexibly applied. We further find that the optimal 3D DLIR configuration is when FBA is combined with a spatial transformer and a DLIR backbone modified with spatial and channel attention, which outperforms existing 3D DLIR approaches. FBA's good performance suggests that spatial attention is a good way to enable scaling up from 2D DLIR to 3D and that a focus on the quality of the image after registration warping is a good way to enhance DLIR performance. Codes and data are available at: https://github.com/kamruleee51/Feedback_DLIR. Md. Kamrul Hasan 0002, Yihao Luo, Guang Yang 0006, Choon Hwai Yap |
IEEE Trans. Medical Imaging | 2 |
| 2025 | Enhanced DTCMR With Cascaded Alignment and Adaptive DiffusionabstractDiffusion tensor cardiovascular magnetic resonance (DTCMR) is the only non-invasive method for visualizing myocardial microstructure, but it is challenged by inconsistent breath-holds and imperfect cardiac triggering, causing in-plane shifts and through-plane warping with an inadequate tensor fitting. While rigid registration corrects in-plane shifts, deformable registration risks distorting the diffusion distribution, and selecting a reference frame among low SNR frames is challenging. Existing pairwise deep learning and iterative methods are unsuitable for DTCMR due to their inability to handle the drastic in-plane motion and disentangle the diffusion contrast distortion with through-plane motions on low SNR frames, which compromises the accuracy of clinical biomarker tensor estimation. Our study introduces a novel deep learning framework incorporating tensor information for groupwise deformable registration, effectively correcting intra-subject inter-frame motion. This framework features a cascaded registration branch for addressing in-plane and through-plane motions and a parallel branch for generating pseudo-frames with diffusion contrasts and template updates to guide registration with a refined loss function and denoising. We evaluated our method on four DTCMR-specific metrics using data from over 900 cases from 2012 to 2023. Our method outperformed three traditional and two deep learning-based methods, achieving reduced fitting errors, the lowest percentage of negative eigenvalues at 0.446%, the highest R2 of HA line profiles at 0.911, no negative Jacobian Determinant, and the shortest reference time of 0.06 seconds per case. In conclusion, our deep learning framework significantly improves DTCMR imaging by effectively correcting inter-frame motion and surpassing existing methods across multiple metrics, demonstrating substantial clinical potential. Fanwen Wang, Yihao Luo, Camila Munoz, Yaqing Luo, Yinzhe Wu 0001, Zohya Khalique, Maria Molto, Ramyah Rajakulasingam, Ranil De Silva, Dudley Pennell, Pedro F. Ferreira, Andrew D. Scott, Sonia Nielles-Vallespin, Guang Yang 0006 |
IEEE Trans. Medical Imaging | 2 |
| 2024 | BootRIST: Detecting and Isolating Mercurial Cores at the Booting Stage
Yihao Luo, Yunjie Deng 0001, Jingquan Ge, Zhenyu Ning, Fengwei Zhang |
ESORICS (2) | 1 |
| 2024 | A Fourier Perspective of Feature Extraction and Adversarial Robustness
Liangqi Zhang, Yihao Luo, Haibo Shen, Tianjiang Wang |
IJCAI | 2 |
| 2024 | Groupwise Deformable Registration of Diffusion Tensor Cardiovascular Magnetic Resonance: Disentangling Diffusion Contrast, Respiratory and Cardiac Motions
Fanwen Wang, Yihao Luo, Pedro F. Ferreira, Yaqing Luo, Yinzhe Wu 0001, Camila Munoz, Dudley Pennell, Andrew D. Scott, Sonia Nielles-Vallespin, Guang Yang 0006 |
MICCAI (2) | 2 |
| 2023 | Training Stronger Spiking Neural Networks with Biomimetic Adaptive Internal Association NeuronsabstractAs the third generation of neural networks, spiking neural networks (SNNs) are dedicated to exploring more insightful neural mechanisms to achieve near-biological intelligence. Intuitively, biomimetic mechanisms are crucial to understanding and improving SNNs. For example, the associative long-term potentiation (ALTP) phenomenon suggests that in addition to learning mechanisms between neurons, there are associative effects within neurons. However, most existing methods only focus on the former and lack exploration of the internal association effects. In this paper, we propose a novel Adaptive Internal Association (AIA) neuron model to establish previously ignored influences within neurons. Consistent with the ALTP phenomenon, the AIA neuron model is adaptive to input stimuli, and internal associative learning occurs only when both dendrites are stimulated at the same time. In addition, we employ weighted weights to measure internal associations and introduce intermediate caches to reduce the volatility of associations. Extensive experiments on prevailing neuromorphic datasets show that the proposed method can potentiate or depress the firing of spikes more specifically, resulting in better performance with fewer spikes. It is worth noting that without adding any parameters at inference, the AIA model achieves state-of-the-art performance on DVS-CIFAR10 (83.9%) and N-CARS (95.64%) datasets. Haibo Shen, Yihao Luo, Liangqi Zhang, Juyu Xiao, Tianjiang Wang |
ICASSP | 2 |
| 2023 | Training Robust Spiking Neural Networks on Neuromorphic Data with Spatiotemporal FragmentsabstractNeuromorphic vision sensors (event cameras) are inherently suitable for spiking neural networks (SNNs) and provide novel neuromorphic vision data for this biomimetic model. Due to the spatiotemporal characteristics, novel data augmentations are required to process the unconventional visual signals of these cameras. In this paper, we propose a novel Event Spatio Temporal Fragments (ESTF) augmentation method. It preserves the continuity of neuromorphic data by drifting or inverting fragments of the spatiotemporal event stream to simulate the disturbance of brightness variations, leading to more robust spiking neural networks. Extensive experiments are performed on prevailing neuromorphic datasets. It turns out that ESTF provides substantial improvements over pure geometric transformations and outperforms other event data augmentation methods. It is worth noting that the SNNs with ESTF achieve the state-of-the-art accuracy of 83.9% on the CIFAR10-DVS dataset. Haibo Shen, Yihao Luo, Liangqi Zhang, Juyu Xiao, Tianjiang Wang |
ICASSP | 2 |
| 2023 | Training Robust Spiking Neural Networks with Viewpoint Transform and Spatiotemporal StretchingabstractNeuromorphic vision sensors (event cameras) simulate biological visual perception systems and have the advantages of high temporal resolution, less data redundancy, low power consumption, and large dynamic range. Since both events and spikes are modeled from neural signals, event cameras are inherently suitable for spiking neural networks (SNNs), which are considered promising models for artificial intelligence (AI) and theoretical neuroscience. However, the unconventional visual signals of these cameras pose a great challenge to the robustness of spiking neural networks. In this paper, we propose a novel data augmentation method, View-Point Transform and SpatioTemporal Stretching (VPT-STS). It improves the robustness of SNNs by transforming the rotation centers and angles in the spatiotemporal domain to generate samples from different viewpoints. Furthermore, we introduce the spatiotemporal stretching to avoid potential information loss in viewpoint transformation. Extensive experiments on prevailing neuromorphic datasets demonstrate that VPT-STS is broadly effective on multi-event representations and significantly outperforms pure spatial geometric transformations. Notably, the SNNs model with VPT-STS achieves a state-of-the-art accuracy of 84.4% on the DVS-CIFAR10 dataset. Haibo Shen, Juyu Xiao, Yihao Luo, Liangqi Zhang, Tianjiang Wang |
ICASSP | 3 |
| 2023 | An Application of Quantum Mechanics to Attention Methods in Computer VisionabstractThis work proposes the quantum-state-based mapping (QSM) for machine learning. QSM uses wave functions that describe microscopic particle systems as mappings. By QSM, original inputs or features extracted by neural networks are processed as quantum states to train wave function parameters. QSM has a low computational cost, almost no additional parameters, and is easy to integrate with other modules. We demonstrate the simplest form of the wave function as a mapping, that is, when a one-dimensional particle is in an infinitely deep potential well, in combination with advanced attention modules. Experiments show that QSM significantly improves the feature recalibration ability of attention module in transfer learning tasks. Then, we tried to analyze the effectiveness of QSM. This work indicates that QSM has an important application value in interdisciplinary machine learning. Yihao Luo, Zehan Li, Wenbo An |
ICASSP | 2 |
| 2023 | Frequency and Scale Perspectives of Feature ExtractionabstractConvolutional neural networks (CNNs) have achieved superior performance but still lack clarity about the nature and properties of feature extraction. In this paper, by analyzing the sensitivity of neural networks to frequencies and scales, we find that neural networks not only have low- and mediumfrequency biases but also prefer different frequency bands for different classes, and the scale of objects influences the preferred frequency bands. These observations lead to the hypothesis that neural networks must learn the ability to extract features at various scales and frequencies. To corroborate this hypothesis, we propose a network architecture based on Gaussian derivatives, which extracts features by constructing scale space and employing partial derivatives as local feature extraction operators to separate high-frequency information. This manually designed method of extracting features from different scales allows our GSSDNets to achieve comparable accuracy with vanilla networks on various datasets. Liangqi Zhang, Yihao Luo, Haibo Shen, Tianjiang Wang |
ICASSP | 2 |
| 2023 | D-IF: Uncertainty-aware Human Digitization via Implicit Distribution FieldabstractRealistic virtual humans play a crucial role in numerous industries, such as metaverse, intelligent healthcare, and self-driving simulation. But creating them on a large scale with high levels of realism remains a challenge. The utilization of deep implicit function sparks a new era of image-based 3D clothed human reconstruction, enabling pixel-aligned shape recovery with fine details. Subsequently, the vast majority of works locate the surface by regressing the deterministic implicit value for each point. However, should all points be treated equally regardless of their proximity to the surface? In this paper, we propose replacing the implicit value with an adaptive uncertainty distribution, to differentiate between points based on their distance to the surface. This simple "value ⇒ distribution" transition yields significant improvements on nearly all the baselines. Furthermore, qualitative results demonstrate that the models trained using our uncertainty distribution loss, can capture more intricate wrinkles, and realistic limbs. Code and models are available for research purposes at github.com/psyai-net/D-IF release. Xueting Yang, Yihao Luo, Yuliang Xiu, Zhaoxin Fan |
ICCV | 2 |
| 2023 | SelfTalk: A Self-Supervised Commutative Training Diagram to Comprehend 3D Talking FacesabstractSpeech-driven 3D face animation technique, extending its applications to various multimedia fields.Previous research has generated promising realistic lip movements and facial expressions from audio signals. However, traditional regression models solely driven by data face several essential problems, such as difficulties in accessing precise labels and domain gaps between different modalities, leading to unsatisfactory results lacking precision and coherence.To enhance the visual accuracy of generated lip movement while reducing the dependence on labeled data, we propose a novel framework SelfTalk, by involving self-supervision in a cross-modals network system to learn 3D talking faces. The framework constructs a network system consisting of three modules: facial animator, speech recognizer, and lip-reading interpreter. The core of SelfTalk is a commutative training diagram that facilitates compatible features exchange among audio, text, and lip shape, enabling our models to learn the intricate connection between these factors. The proposed framework leverages the knowledge learned from the lip-reading interpreter to generate more plausible lip shapes. Extensive experiments and user studies demonstrate that our proposed approach achieves state-of-the-art performance both qualitatively and quantitatively. We recommend watching the supplementary video. Ziqiao Peng, Yihao Luo, Xiangyu Zhu 0001, Hongyan Liu 0002, Jun He 0008, Zhaoxin Fan |
ACM Multimedia | 2 |
| 2023 | Dynamic multi-scale loss optimization for object detection
Yihao Luo, Tianjiang Wang, Qi Feng 0003 |
Multim. Tools Appl. | 1 |
| 2022 | Kernel Estimation Network for Blind Super-ResolutionabstractExisting super-resolution (SR) methods commonly assume that the degradation kernels are fixed and known (e.g., bicubic downsampling or single Gaussian blurring kernel). However, these methods suffer a severe performance drop when the real degradations deviate from this assumption. To address this issue, this paper proposes a novel kernel estimation network (KENet) for kernel prediction. Specifically, KENet predicts the degradation kernels by optimizing the kernel space loss in a supervised way, without extra iterations at the inference time. Moreover, we introduce an adaptive attention loss to constrain the kernel optimization space, which can bias the allocation of trainable model parameters towards the most informative components of the estimation kernels. Extensive experiments on synthetic and real images show that the proposed KENet not only encourages a more accurate way to predict degradation kernels but also outperforms existing state-of-the-art blind SR methods when combined with non-blind SR methods. Haibo Shen, Liangqi Zhang, Yihao Luo, Tianjiang Wang |
ICASSP | 4 |
| 2022 | Multi-View Data Representation Via Deep Autoencoder-Like Nonnegative Matrix FactorizationabstractSince a large proportion of real-world data is made of different representations or views, learning on data represented with multiple views (e.g., numerous types of features or modalities) has garnered considerable attention recently. Nonnegative matrix factorization (NMF) has been widely adopted for multi-view learning due to its great interpretability. We focus on unsupervised multi-view data representation in this paper and propose a novel framework termed Deep Autoencoder-like NMF (DANMF-MDR), which learns an intact representation by simultaneously exploring multi-view complementary and consistent information. Furthermore, an efficient iterative optimization algorithm is developed to solve the proposed model. Experimental results on three real-world multi-view datasets demonstrate that ours performs better than the SOTA multi-view NMF-based MDR approaches. Haonan Huang, Yihao Luo, Guoxu Zhou, Qibin Zhao |
ICASSP | 2 |
| 2022 | Dynamic Multi-Scale Loss Balance for Object DetectionabstractIt is a common paradigm in object detection frameworks to perform multi-scale detection. However, each scale is treated equally during training. In this paper, we carefully study the objective imbalance of multi-scale detector training. We argue that the loss in each scale is neither equally important nor independent. Different from the existing solutions of setting fixed multi-task weights, we dynamically optimize the loss weight of each scale in the training process. Specifically, we propose an Adaptive Variance Weighting (AVW) to balance multi-scale loss according to the statistical variance. Then we develop a novel Reinforcement Learning Optimization (RLO) to decide the weighting scheme probabilistically during training. The proposed dynamic methods make better utilization of multi-scale training loss without extra computational complexity and learnable parameters for backpropagation. Experiments on Pascal VOC and MS COCO benchmark validate the effectiveness of our proposed methods. Yihao Luo, Tianjiang Wang, Qi Feng 0003 |
ICASSP | 1 |
| 2022 | Multi-Scale Reinforcement Learning Strategy for Object DetectionabstractFeature Pyramid Network (FPN) has become a common detection paradigm by improving multi-scale features with strong semantics. However, most FPN-based methods typically treat each feature map equally and sum the loss without distinction, which might lead to suboptimal overall performance. In this paper, we propose a Multi-scale Reinforcement Learning Strategy (MRLS) for balanced multi-scale training. First, we design Dynamic Feature Fusion (DFF) to dynamically magnify the impact of more important feature maps in FPN. Second, we introduce Compensatory Scale Training (CST) to enhance the supervision of the under-training scale. We regard the whole detector as a reinforcement learning system while the state bases on multi-scale loss. And we develop the corresponding action, reward, and policy. Compared with adding more rich model architectures, MRLS would not add any extra modules and computational burdens on the baselines. Experiments on MS COCO and PASCAL VOC benchmark demonstrate that our method significantly improves the performance of commonly used object detectors. Yihao Luo, Leixilan Pan, Tianjiang Wang, Qi Feng 0003 |
ICASSP | 1 |
| 2022 | Efficient CNN Architecture Design Guided by VisualizationabstractModern efficient Convolutional Neural Networks(CNNs) always use Depthwise Separable Convolutions(DSCs) and Neural Architecture Search(NAS) to reduce the number of parameters and the computational complexity. But some inherent characteristics of networks are overlooked. Inspired by visualizing feature maps and N×N(N>1) convolution kernels, several guidelines are introduced in this paper to further improve parameter efficiency and inference speed. Based on these guidelines, our parameter-efficient CNN architecture, called VGNetG, achieves better accuracy and lower latency than previous networks with about 30%~50% parameters reduction. Our VGNetG-1.0MP achieves 67.7% top-1 accuracy with 0.99M parameters and 69.2% top-1 accuracy with 1.14M parameters on ImageNet classification dataset. Furthermore, we demonstrate that edge detectors can replace learnable depthwise convolution layers to mix features by replacing the N×N kernels with fixed edge detection ker-nels. And our VGNetF-1.5MP archives 64.4%(-3.2%) top-1 accuracy and 66.2%(-1.4%) top-1 accuracy with additional Gaussian kernels. Liangqi Zhang, Haibo Shen, Yihao Luo, Leixilan Pan, Tianjiang Wang, Qi Feng 0003 |
ICME | 3 |
| 2022 | Dual-branch network via pseudo-label training for thyroid nodule detection in ultrasound image
Ruoning Song, Chuang Zhu, Long Zhang 0020, Yihao Luo, Jun Liu 0014, Jie Yang 0023 |
Appl. Intell. | 5 |
| 2022 | Contour information regularized tensor ring completion for realistic image restorationabstractAbstract Tensor completion has gained considerable research interest in recent years and has been frequently applied to image restoration. This type of method basically employs the low‐rank nature of images, implicitly requiring that the whole picture is of globally consistent features. As a result, existing tensor completion algorithms often give reasonably good performance if the target image has only random pixel‐level missing. Unfortunately, pixel‐level missing is very rare in practice and it is often wanted to restore an image with irregular hole‐shaped missing, such as removing electricity poles from landscape photos or irrelevant people from tourist photos. This task is extremely difficult for traditional low‐rank based tensor completion methods. To overcome this drawback, a Contour Information regularized Tensor RIng Completion (CITRIC) method is proposed for practical image restoration. Meanwhile, the contour information regularization is used to capture significant local features, whereas the low‐rank tensor ring structure is utilized to capture as much global information as possible. The alternating direction method of multipliers (ADMM) is adopted to optimize the cost function. Extensive experimental results using real‐world images show that CITRIC is more practical than existing methods and can restore real‐world images with irregular hole‐shaped missing. Yihao Luo, Zhifa Liu, Guoxu Zhou |
IET Image Process. | 2 |
| 2022 | CE-FPN: enhancing channel information for object detection
Yihao Luo, Jingjuan Guo, Haibo Shen, Tianjiang Wang, Qi Feng 0003 |
Multim. Tools Appl. | 1 |
| 2022 | Conversion of Siamese networks to spiking neural networks for energy-efficient object tracking
Yihao Luo, Haibo Shen, Tianjiang Wang, Qi Feng 0003, Zehan Tan |
Neural Comput. Appl. | 1 |
| 2021 | SiamSNN: Siamese Spiking Neural Networks for Energy-Efficient Object Tracking
Yihao Luo, Caihong Yuan, Liangqi Zhang, Tianjiang Wang, Qi Feng 0003 |
ICANN (5) | 1 |
| 2021 | SFCN: Symmetric feature comparison network for detecting ischemic stroke lesions on CT imagesabstractAbstract Ischemic stroke is the most common stroke and the leading cause of disability and death in the world. Computed tomography (CT) is a popular and economical diagnostic device for the stroke, However the ischemic stroke lesions are not evident on CT images and the diagnostic result relies on the visual observation of neurologists, which may vary from doctor to doctor. To facilitate the treatment, a computer‐aided detection algorithm on CT images is proposed to help clinician for the ischemic stroke screening. In order to obtain accurate lesion annotation on CT images, novel automatic algorithms are developed to achieve image pairing, calibration, and registration. Then, a new framework with the symmetric feature extraction and comparison is proposed to identify and locate the ischemic stroke lesion. Experimental results show that this method achieves 75% of DICE in the detection of ischemic stroke lesions, which is higher than other methods by 4%. Its competitive results compared with seven latest methods is shown in terms of extensive qualitative and quantitative evaluation. This method can accurately detect the lesion in the CT images through the comparison of symmetric regional features, which has contributed to the clinical diagnosis of ischemic stroke. Long Zhang 0020, Chuang Zhu, Yuewei Wu, Yang Yang 0007, Yihao Luo, Ruoning Song, Jie Yang 0023 |
IET Image Process. | 5 |
| 2021 | Blind image super-resolution based on prior correction network
Yihao Luo, Yi Xiao 0004, Xianyi Zhu, Tianjiang Wang, Qi Feng 0003, Zehan Tan |
Neurocomputing | 2 |
| 2021 | DAEANet: Dual auto-encoder attention network for depth map super-resolution
Yihao Luo, Xianyi Zhu, Liangqi Zhang, Haibo Shen, Tianjiang Wang, Qi Feng 0003 |
Neurocomputing | 2 |
| 2021 | Multi-level colonoscopy malignant tissue detection with adversarial CAC-UNet
Chuang Zhu, Ke Mei, Yihao Luo, Jun Liu 0014, Ying Wang 0043, Mulan Jin |
Neurocomputing | 4 |
| 2020 | Object detector with enriched global context information
Jingjuan Guo, Caihong Yuan, Ping Feng, Yihao Luo, Tianjiang Wang |
Multim. Tools Appl. | 5 |
| 2019 | Breast Cancer Image Classification on WSI with Spatial CorrelationsabstractAs common cancer, breast cancer kills thousands of women every year. It’s significant to provide doctors computer-aided diagnosis (CAD) to ease their workload as well as improve detection quality. Patch-level CNNs are usually used to classify the breast tissue slice, and the CNNs classify each patch independently ignoring the spatial correlations, resulting in wrong isolated label map. However, the probability distribution of cancer type is related to their adjacent patches. In this paper, we propose a framework integrating CNN and filter algorithm aimed at extracting spatial information and improving the performance of the classification. The network was trained on a breast cancer dataset provided by ICIAR18. For 4-class classification, compared to CNN methods without using spatial correlations, the proposed method achieved about 10% improvement on accuracy over the validation dataset and get smoother probability maps. Our experiments also show that larger kernel size gets better performance. The code is available at https://github.com/dong100136/Breast-Cancer-Image-Classification-On-WSI-With-Spatial-Correlations. Jiandong Ye, Yihao Luo, Chuang Zhu, Yue Zhang 0016 |
ICASSP | 2 |
| 2019 | A Spiking Neural Network Architecture for Object Tracking
Yihao Luo, Quanzheng Yi, Tianjiang Wang, Caihong Yuan, Jingjuan Guo, Ping Feng, Qi Feng 0003 |
ICIG (1) | 1 |
| 2019 | Learning deep embedding with mini-cluster loss for person re-identification
Caihong Yuan, Jingjuan Guo, Ping Feng, Yihao Luo, Chunyan Xu, Tianjiang Wang, Kui Duan |
Multim. Tools Appl. | 5 |