Lin Gu 0003

dblp:70/3413-3 · DBLP profile ↗
← Back
85ranked-venue papers
10as first author
64since 2021 · last 2026
0000-0002-7419-6240ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 58 · 4 first-author · 47 since 2021Graphics, computer vision, multimedia, augmented reality and games · 52 · 7 first-author · 36 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Perception-Inspired Color Space Design for Photo White Balance Editing
abstract
White balance (WB) is a key step in the image signal processor (ISP) pipeline that mitigates color casts caused by varying illumination and restores the scene’s true colors. Currently, sRGB-based WB editing for post-ISP WB correction is widely used [2], [18] to address color constancy failures in the ISP pipeline when the original camera RAW is unavailable. However, additive color models (e.g., sRGB) are inherently limited by fixed nonlinear transformations and entangled color channels, which often impede their generalization to complex lighting conditions.To address these challenges, we propose a novel framework for WB correction that leverages a perception-inspired Learnable HSI (LHSI) color space. Built upon a cylindrical color model that naturally separates luminance from chromatic components, our framework further introduces dedicated parameters to enhance this disentanglement and learnable mapping to adaptively refine the flexibility. Moreover, a new Mamba-based network is introduced, which is tailored to the characteristics of the proposed LHSI color space.Experimental results on benchmark datasets demonstrate the superiority of our method, highlighting the potential of perception-inspired color space design in computational photography. The source code is avail-able at https://github.com/YangCheng58/WB_Color_Space.
Yang Cheng 0008, Ziteng Cui, Lin Gu 0003, Shenghan Su, Zenghui Zhang
WACV3
2026 DNGaussian++: Improving Sparse-View Gaussian Radiance Fields With Depth Normalization
abstract
Synthesizing novel views from sparse views has achieved impressive advances with radiance fields, yet prevailing methods suffer from high consumption or insufficient refinement capability. This paper introduces DNGaussian, a depth-regularized framework based on 3D Gaussian Splatting, offering real-time and high-quality few-shot novel view synthesis at low costs. Our motivation stems from the remarkable advancement of recent 3D Gaussian Splatting, despite it will encounter a geometry degradation when input views decrease. In the Gaussian radiance fields, we find this degradation in scene geometry primarily lined to the positioning of Gaussian primitives and can be mitigated by depth constraint. Consequently, we propose a Hard and Soft Depth Regularization to restore accurate scene geometry under coarse monocular depth supervision while maintaining a fine-grained color appearance. To further refine detailed geometry, we introduce Global-Local Depth Normalization, enhancing the focus on small local depth changes. Although DNGaussian shows impressive performance, its patch-wise regularization obscures the inconsistency in cross-patch errors. Additionally, primitives can still be irreversibly trapped in local minima under sparse views, even if depth regularization is applied. In this paper, we propose an extended version, DNGaussian++. First, a Geometry Instance Regularizer is developed to enable depth regularization for continuous consistency by exploiting reliable instance-level depth cues. Leveraging the depth gradient guidance, we then propose a Depth-Guided Geometry Reorganization to address the aforementioned local minima problem with high representation efficiency. Extensive experiments show that DNGaussian++ exhibits state-of-the-art performance in multiple datasets and scenarios with high efficiency, and the broad applicability and effectiveness are verified on various backbones and tasks.
Jiahe Li 0007, Xiaohan Yu 0001, Xiao Bai 0001, Xin Ning 0001, Lin Gu 0003
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 Gradient-Refined Federated Learning on Head-Tail Imbalanced Data
abstract
Federated learning has emerged as a transformative paradigm for distributed data collaboration, facilitating knowledge aggregation across multiple local clients through a global server while rigorously preserving data privacy. However, its performance is significantly hindered by the global head-tail imbalance, where tail classes with scarce data are often dominated by head classes. This challenge, known as federated long-tailed learning, arises from the intrinsic conflict between class knowledge acquisition and privacy preservation. Existing methodologies falter in resolving this conflict, as the abstraction of data knowledge in federated communication complicates the extraction of class-level knowledge, resulting in imbalanced global models and diminished performance. To simultaneously address this imbalance and uphold privacy, we introduce FedGRE, a gradient-refined federated learning approach that constructs global gradients and facilitates refined global gradient descent. FedGRE enhances gradients through two pivotal mechanisms: accumulation diffusion and accumulation refinement. The former amalgamates accumulated gradients with stochastic gradient perturbations to alleviate class imbalance, while the latter utilizes the accumulation as an anchor to calibrate global gradient updates, ensuring consistency and mitigating oscillations. Additionally, we implement a consistency integration technique to incorporate the refined accumulation into the global model, guaranteeing privacy-preserving and class-balanced global optimization. Extensive experiments on six datasets demonstrate that FedGRE significantly outperforms 14 state-of-the-art (SOTA) methods in federated long-tailed classification while maintaining robust privacy protection.
Heye Zhang, Chenchu Xu, Lin Gu 0003, Jingfeng Zhang, Tieyong Zeng, Zhifan Gao
IEEE Trans. Neural Networks Learn. Syst.4
2025 EventPillars: Pillar-based Efficient Representations for Event Data
abstract
Event Cameras offer appealing advantages, including power efficiency and ultra-low latency, driving forward advancements in edge applications. In order to leverage mature frame-based algorithms, most approaches typically compute dense, image-like representations from sparse, asynchronous events. However, they are often unable to capture comprehensive information or are computationally intensive, which hinders the edge deployment of event-based vision. Meanwhile, pillar-based paradigms have been proven to be efficient and well established for dense representations of sparse data. Hence, from a novel pillar-based perspective, we present EventPillars, an efficient, comprehensive framework for dense event representations. To summarize, it (i) incorporates the Temporal Event Range to describe an intact temporal distribution, (ii) Activates the Event Polarities to explicitly record the scene dynamics, (iii) enhances the target awareness by a spatial attention prior from Normalized Event Density, (iv) can be plug-and-played into different downstream tasks. Extensive experiments show that our EventPillars records a new state-of-the-art precision on object recognition and detection datasets with surprisingly 9.2× and 4.5× lower computation and storage consumption. This brings a new insight into dense event representations and is promising to boost the edge deployment of event-based vision.
Rui Fan 0001, Weidong Hao, Juntao Guan, Lai Rui, Lin Gu 0003, Fanhong Zeng, Zhangming Zhu
AAAI5
2025 TdAttenMix: Top-Down Attention Guided Mixup
abstract
CutMix is a data augmentation strategy that cuts and pastes image patches to mixup training data. Existing methods pick either random or salient areas which are often inconsistent to labels, thus misguiding the training model. By our knowledge, we integrate human gaze to guide cutmix for the first time. Since human attention is driven by both high-level recognition and low-level clues, we propose a controllable Top-down Attention Guided Module to obtain a general artificial attention which balances top-down and bottom-up attention. The proposed TdATttenMix then picks the patches and adjust the label mixing ratio that focuses on regions relevant to the current label. Experimental results demonstrate that our TdAttenMix outperforms existing state-of-the-art mixup methods across eight different benchmarks. Additionally, we introduce a new metric based on the human gaze and use this metric to investigate the issue of image-label inconsistency.
Lin Gu 0003, Feng Lu 0001
AAAI2
2025 InsTaG: Learning Personalized 3D Talking Head from Few-Second Video
abstract
Despite exhibiting impressive performance in synthesizing lifelike personalized 3D talking heads, prevailing methods based on radiance fields suffer from high demands for training data and time for each new identity. This paper introduces InsTaG, a 3D talking head synthesis framework that allows a fast learning of realistic personalized 3D talking head from few training data. Built upon a lightweight 3DGS person-specific synthesizer with universal motion priors, InsTaG achieves high-quality and fast adaptation while preserving high-level personalization and efficiency. As preparation, we first propose an Identity-Free Pre-training strategy that enables the pre-training of the person-specific model and encourages the collection of universal motion priors from long-video data corpus. To fully exploit the universal motion priors to learn an unseen new identity, we then present a Motion-Aligned Adaptation strategy to adaptively align the target head to the pre-trained field, and constrain a robust dynamic head structure under few training data. Experiments demonstrate our outstanding performance and efficiency under various data scenarios to render high-quality personalized talking heads. Project page: https://fictionarry.github.io/InsTaG/.
Jiahe Li 0007, Xiao Bai 0001, Jun Zhou 0001, Lin Gu 0003
CVPR6
2025 Frequency Dynamic Convolution for Dense Image Prediction
abstract
While Dynamic Convolution (DY-Conv) has shown promising performance by enabling adaptive weight selection through multiple parallel weights combined with an attention mechanism, the frequency response of these weights tends to exhibit high similarity, resulting in high parameter costs but limited adaptability. In this work, we introduce Frequency Dynamic Convolution (FDConv), a novel approach that mitigates these limitations by learning a fixed parameter budget in the Fourier domain. FDConv divides this budget into frequency-based groups with disjoint Fourier indices, enabling the construction of frequency-diverse weights without increasing the parameter cost. To further enhance adaptability, we propose Kernel Spatial Modulation (KSM) and Frequency Band Modulation (FBM). KSM dynamically adjusts the frequency response of each filter at the spatial level, while FBM decomposes weights into distinct frequency bands in the frequency domain and modulates them dynamically based on local content. Extensive experiments on object detection, segmentation, and classification validate the effectiveness of FD-Conv. We demonstrate that when applied to ResNet-50, FDConv achieves superior performance with a modest increase of +3.6M parameters, outperforming previous methods that require substantial increases in parameter budgets (e.g., CondConv +90M, KW +76.5M). Moreover, FD-Conv seamlessly integrates into a variety of architectures, including ConvNeXt, Swin-Transformer, offering a flexible and efficient solution for modern vision tasks. The code is made publicly available at https://github.com/Linwei-Chen/FDConv.
Lin Gu 0003, Liang Li 0003, Chenggang Yan 0001, Ying Fu 0001
CVPR2
2025 Frequency-Dynamic Attention Modulation for Dense Prediction
abstract
Vision Transformers (ViTs) have significantly advanced computer vision, demonstrating strong performance across various tasks. However, the attention mechanism in ViTs makes each layer function as a low-pass filter, and the stacked-layer architecture in existing transformers suffers from frequency vanishing. This leads to the loss of critical details and textures. We propose a novel, circuit-theory-inspired strategy called Frequency-Dynamic Attention Modulation (FDAM), which can be easily plugged into ViTs. FDAM directly modulates the overall frequency response of ViTs and consists of two techniques: Attention Inversion (AttInv) and Frequency Dynamic Scaling (FreqScale). Since circuit theory uses low-pass filters as fundamental elements, we introduce AttInv, a method that generates complementary high-pass filtering by inverting the low-pass filter in the attention matrix, and dynamically combining the two. We further design FreqScale to weight different frequency components for fine-grained adjustments to the target response function. Through feature similarity analysis and effective rank evaluation, we demonstrate that our approach avoids representation collapse, leading to consistent performance improvements across various models, including SegFormer, DeiT, and MaskDINO. These improvements are evident in tasks such as semantic segmentation, object detection, and instance segmentation. Additionally, we apply our method to remote sensing detection, achieving state-of-the-art results in single-scale settings. The code is available at https://github.com/Linwei-Chen/FDAM.
Lin Gu 0003, Ying Fu 0001
ICCV2
2025 GeoSVR: Taming Sparse Voxels for Geometrically Accurate Surface Reconstruction
abstract
Reconstructing accurate surfaces with radiance fields has achieved remarkable progress in recent years. However, prevailing approaches, primarily based on Gaussian Splatting, are increasingly constrained by representational bottlenecks. In this paper, we introduce GeoSVR, an explicit voxel-based framework that explores and extends the under-investigated potential of sparse voxels for achieving accurate, detailed, and complete surface reconstruction. As strengths, sparse voxels support preserving the coverage completeness and geometric clarity, while corresponding challenges also arise from absent scene constraints and locality in surface refinement. To ensure correct scene convergence, we first propose a Voxel-Uncertainty Depth Constraint that maximizes the effect of monocular depth cues while presenting a voxel-oriented uncertainty to avoid quality degradation, enabling effective and robust scene constraints yet preserving highly accurate geometries. Subsequently, Sparse Voxel Surface Regularization is designed to enhance geometric consistency for tiny voxels and facilitate the voxel-based formation of sharp and accurate surfaces. Extensive experiments demonstrate our superior performance compared to existing methods across diverse challenging scenarios, excelling in geometric accuracy, detail preservation, and reconstruction completeness while maintaining high efficiency. Code is available at https://github.com/Fictionarry/GeoSVR.
Jiahe Li 0007, Youmin Zhang 0005, Xiao Bai 0001, Xiaohan Yu 0001, Lin Gu 0003
NeurIPS7
2025 I2-NeRF: Learning Neural Radiance Fields Under Physically-Grounded Media Interactions
abstract
Participating in efforts to endow generative AI with the 3D physical world perception, we propose I2-NeRF, a novel neural radiance field framework that enhances isometric and isotropic metric perception under media degradation. While existing NeRF models predominantly rely on object-centric sampling, I2-NeRF introduces a reverse-stratified upsampling strategy to achieve near-uniform sampling across 3D space, thereby preserving isometry. We further present a general radiative formulation for media degradation that unifies emission, absorption, and scattering into a particle model governed by the Beer–Lambert attenuation law. By matting direct and media-induced in-scatter radiance, this formulation extends naturally to complex media environments such as underwater, haze, and even low-light scenes. By treating light propagation uniformly in both vertical and horizontal directions, I2-NeRF enables isotropic metric perception and can even estimate medium properties such as water depth. Experiments on real-world datasets demonstrate that our method significantly improves both reconstruction fidelity and physical plausibility compared to existing approaches. The source code is available at https://github.com/ShuhongLL/I2-NeRF.
Shuhong Liu, Lin Gu 0003, Ziteng Cui, Xuangeng Chu, Tatsuya Harada
NeurIPS2
2025 Physiology-Aware PolySnake for Coronary Vessel Segmentation
abstract
Coronary artery disease (CAD) is a significant health risk that requires early detection for effective treatment. While recent advances in deep learning have shown promise in automating CAD detection from coronary computed to-mography angiography (CCTA) images, the accurate segmentation of coronary vessels remains a challenge, particularly due to the imbalanced presence of plaque in unhealthy vessels. This paper introduces a physiology-aware approach11https://github.com/opensourcetorch/Physiology-aware-PolySnake to coronary vessel segmentation that addresses these challenges. Our proposed pipeline consists of three main components. First, a hybrid UNeXt architecture is designed to segment artery boundaries and predict initial boundary contours by leveraging 3D spatial relations among adjacent slices. Second, we introduce multi-class circular convolution for iterative contour deformation, which generates well-connected contour pairs of the artery wall's inner and outer boundaries through iterative refinement. Finally, we propose a focal smooth Lllossfunction to handle the implicit class imbalance caused by plaque in unhealthy vessels and to enhance the robustness of the physiology-aware polysnake network by explicitly limiting the accuracy of initial contours. Extensive evaluations demonstrate that our methods significantly improve model performance, achieving state-of-the-art results in coronary vessel segmentation.
Yizhe Ruan, Lin Gu 0003, Yusuke Kurose, Junichi Iho, Youji Tokunaga, Makoto Horie, Yusaku Hayashi, Keisuke Nishizawa, Yasushi Koyama, Tatsuya Harada
WACV2
2025 Scene-aware contrastive regression for multi-person action quality assessment
Xini Ding, Chun-Ting Wang, Lin Gu 0003
Appl. Intell.6
2025 Guest Editorial: Multi-view representation learning for computer vision
abstract
Object recognition and scene analysis in single-view images may face difficulties such as occlusion and incomplete information, while multi-view learning can address this limitation. When an object or scene is observed from multiple views, information on target objects can be significantly enriched to improve the performance of computer vision tasks. For this reason, multi-view has become one of the important forms of data representation, which leads to the emerging of new research topics on complete or in-complete multi-view learning. Multi-view learning enables the use of multi-source information, nevertheless, the heterogeneous characteristics of data make it difficult to reliably associate information from different views, especially in a complex environment. It remains challenging for tasks to make effective use of the consistent and complementary information between different complete views and to enhance the completeness of potential representation. A wide variety of research is being conducted to explore and discover possible challenges and opportunities to exploit multi-view representation learning for computer vision. The purpose of this Special Issue is to collect high-quality articles on the recent development and trend of multi-view representation learning in computer vision, publish new ideas, theories, solutions and insights on this topic, and showcase their applications. In this Special Issue, we have received 36 papers, all of which underwent peer review. Of the 36 originally submitted papers, 10 have been accepted, which cover a variety of fields, such as person re-identification, gait recognition, 3D object recognition, and behaviour recognition. These accepted papers are mainly divided into three categories. The first category covers the incomplete multi-view data learning theoretics and methods. The papers in this category are of He et al., Kun et al., Fan et al. and Wang et al. The last two categories are both multi-view applications. One of which is 3D-related applications. The papers in this category are of Qi et al. and Sun et al. The other category is about 2D recognition. The papers in this category are of Zhang et al., Huang et al., Zheng et al. and Zhang et al. A brief presentation of each of the paper follows. He et al. present an innovative multi-view subspace clustering method with incomplete graph information. Specifically, they separate one shared and multiple specific graphs from multiple raw graph data, and exploit the mask fusion strategy and block diagonal regulariser to obtain the inherent category information. The clustering results on six real-world datasets show that the method outperforms a series of classic incomplete multi-view clustering methods. Kun et al. propose a new method for low-rank-based multi-view subspace clustering based on low-rank correlation analysis. To overcome the limitations of unreliable low-rank structure and imprecise graphs caused by multi-view noise and outliers, they introduce the canonical correlation analysis strategy and a dual regularisation term to characterise the connections between different views adaptively. Experimental results reveal the method's superiority over compared state-of-the-art (SOTA) methods in accuracy, normalised mutual information, and F-score evaluation metrics. Fan et al. address the challenge of partial mapping between the views in multi-view clustering, and propose a self-inferring incomplete multi-view clustering algorithm to explore the information hidden in the local geometric structure and recover missing instances through mining the information hidden in existing instances. Experimental results show that the method can improve the clustering performance compared with the SOTA methods. Wang et al. propose a semi-paired semi-supervised deep hashing to solve the large-scale multimedia retrieval task. The method is an end-to-end deep neural network model with high-order affinity. To maintain the consistency within the modalities, they introduce a common representation that combines with the labelled information to associate different modalities. Experimental results demonstrate the superior performance of proposed method. Qi et al. propose a double-weighting convolution neural network based on the L2-S grouping mechanism for multi-view 3D object recognition. The goal of the proposed L2-S grouping mechanism is to calculate the discrimination score of views and group views more reasonably. Results of the experiments show that the method can achieve SOTA performance. Sun et al. present a dual-matching method with cross-attention mechanism to address the limitations of matching-based methods caused by a preset fixed disparity range on depth estimation task. To tackle the mismatches on edges and details, they introduce an exquisite module based on left-right consistency. The method is proved to be competitive and effective by experiments conducted under popular benchmarks. Zhang et al. want to answer the following two questions: (1) does a query image with higher resolution than that of the gallery image also affect the pedestrian re-identification performance? If so, and (2) how does it affect performance? So, they propose an end-to-end trainable resolution independent person re-identification network that is composed of a cross-resolution Generative Adversarial Networks and embedding batch normalisation layers. The results demonstrate that the proposed method outperforms the SOTA methods in the pedestrian re-identification task on their expanded benchmark dataset. Huang et al. address the limitation of current gait-based age and gender recognition methods under multi-view scene, and propose an attention-aware spatio–temporal learning framework that employs silhouette sequence as an input to learn essential spatial–temporal gait representation. The proposed method has produced results that outperformed the benchmarks with an Mean Absolute Error of 6.68 years for age estimation and a Correct Classification Rate of 97% for gender classification. Zheng et al. apply deep learning to multi-view classroom behaviour detection. First, they propose an improved detection model based on YOLOv5 to improve the convergence speed of the prediction box. Second, they establish a quantitative evaluation standard for students' classroom attention, and then conduct training and verification by collecting multi-view classroom datasets. Finally, they increase the environment variation in the training model phase to make the model have better generalisation ability. Experiments demonstrate that the method can effectively identify and detect students' behaviours in the classroom from different views. Zhang et al. propose a method for multi-dimensional video anomaly detection, which uses the Object-meta instead of video frames as the input, and the Memory Search Guided Autoencoder with Memory Pools (MSGAE-MP) to reconstruct. The multi-dimensional information carried by the input can be strengthened via Object-meta. The MSGAE-MP construct multi-level memory pools, so as to reconstruct Object-meta in different dimensions. Experiments show that the method is feasible and has achieved excellent results. All of the papers published in this Special Issue show that multi-view representation learning theoretics have developed very fast in recent years. In addition, it is very promising to solve traditional computer vision tasks under multi-view setting, including but not limited to 3D object recognition, person re-identification, gait-based age and gender estimation, and depth estimation. Xin Ning and Chen Wang are responsible for the writing of Proposal and Editorial materials; Jun Zhou is responsible for the processing of articles; and Jing Wu, Lin Gu and Jian Cheng are responsible for the solicitation and publicity of the special issue. Firstly, we would like to thank all the authors for their innovative contributions and all the reviewers for their professional and crucial, yet constructive comments. Also, we wish to express our thanks to Mr Hang Ran, PhD students at Institute of Semiconductors, Chinese Academy of Sciences, for his assistance in this process. Last, we wish to express our gratitude to the editorial team of IET Computer Vision for their support throughout this venture. We hope you enjoy this collection of papers and that the Special Issue can stimulate further research and development in this area. This work is supported by the National Natural Science Foundation of China (Grant no. 61901436). National Natural Science Foundation of China, Grant/Award Number: 61901436. Data sharing is not applicable to this article as no new data were created or analyzed in this study. Xin Ning (SMIEEE) received a B.S. degree in software engineering in 2012, and a Ph.D. degree in electronic circuit and system from the university of Chinese Academy of Sciences, in 2017. He is currently an associate professor with the Laboratory of Artificial Neural Networks and High Speed Circuits, Institute of Semiconductors, Chinese Academy of Sciences. His current research interests include neural networks, intelligent systems and computer vision. He has published as the first or corresponding author in more than 45 papers in journals and refereed conferences. Now he serves as the young associated editor of CAAI Transactions on Intelligent Systems, the guest editor of Elsevier Journal on DISPLAYS. He is also the guest editor of CONNECTION SCIENCE and CONCURR COMP-PRACT E. He was the Website Chair of the IEEE HPBD&IS 2020 and the Publication Chair of the IEEE HPBD&IS 2021. Jun Zhou received a B.S. degree in computer science and a B.E. degree in international business from the Nanjing University of Science and Technology, Nanjing, China, in 1996 and 1998, respectively, an M.S. degree in computer science from Concordia University, Montreal, QC, Canada, in 2002, and a Ph.D. degree in computing science from the University of Alberta, Edmonton, AB, Canada, in 2006. He was a research fellow with the Research School of Computer Science, The Australian National University, Canberra, ACT, Australia, and a researcher with the Canberra Research Laboratory, National Information and Communications Technology Australia, Canberra. In 2012, he joined the School of Information and Communication Technology, Griffith University, Nathan, QLD, Australia, where he is currently a reader. His research interests include pattern recognition, computer vision, and spectral imaging and their applications in remote sensing and environmental informatics. He is the associate editor for the journal of Pattern Recognition and IEEE Trans. on Remote Sensing. Jian Cheng is a professor of Institute of Automation, Chinese Academy of Sciences. He received the B.S. and M.S. degrees in Mathematics from Wuhan University in 1998 and 2001, respectively. After that, he received a Ph.D degree in pattern recognition and intelligent systems from Institute of Automation, Chinese Academy of Sciences in 2004. His current major research interests include deep learning, computer vision, chip design, etc. Jing Wu is now a postdoc at the school of computer science, Beihang University. He received his B.E. degree from the school of computer science, Northwestern Polytechnical University in 2013 and received his PhD. degree from the school of computer science, Beihang University in 2021. His research interests include computer vision, stereo matching, 3D reconstruction and camera localization. Chen Wang is now a postdoc at the school of computer science, Beihang University. He received his B.E. degree from the school of computer science, Northwestern Polytechnical University in 2013 and received his PhD. degree from the school of computer science, Beihang University in 2021. His research interests include computer vision, stereo matching, 3D reconstruction and camera localization. Lin Gu received a B.Eng. degree from Shanghai University, Shanghai, China, in 2009, and a Ph.D. degree in computer vision from Australian National University in 2014. After Ph.D. graduation from the Australian National University, he worked as a post-doctoral researcher at A*STAR, Singapore. Then, he was a project researcher with the National Institute of Informatics, Japan, and also a visiting scholar with Kyoto University, Japan. He is currently a research scientist at RIKEN AIP, Japan, and a special researcher with the University of Tokyo, Japan. He is also an in-charge of a Moonshot and an ACT-X Project to improve artificial intelligence by simulating the human brain. His primary research interests lie in machine learning, medical imaging, and computational photography.
Xin Ning 0001, Jun Zhou 0001, Jian Cheng 0001, Jing Wu 0004, Chen Wang 0026, Lin Gu 0003
IET Comput. Vis.6
2025 Defender of privacy and fairness: Tiny but reversible generative model via mutually collaborative knowledge distillation
Sissi Xiaoxiao Wu, Zehong Huang, Zhicong Liang, Lin Gu 0003, Tatsuya Harada, Yingying Zhu 0004
Neurocomputing4
2025 Adaptive frequency-aware network for action quality assessment
Chun-Ting Wang, Xini Ding, Lin Gu 0003
Multim. Syst.5
2025 Spatial Frequency Modulation for Semantic Segmentation
abstract
High spatial frequency information, including fine details like textures, significantly contributes to the accuracy of semantic segmentation. However, according to the Nyquist-Shannon Sampling Theorem, high-frequency components are vulnerable to aliasing or distortion when propagating through downsampling layers such as strided-convolution. Here, we propose a novel Spatial Frequency Modulation (SFM) that modulates high-frequency features to a lower frequency before downsampling and then demodulates them back during upsampling. Specifically, we implement modulation through adaptive resampling (ARS) and design a lightweight add-on that can densely sample the high-frequency areas to scale up the signal, thereby lowering its frequency in accordance with the Frequency Scaling Property. We also propose Multi-Scale Adaptive Upsampling (MSAU) to demodulate the modulated feature and recover high-frequency information through non-uniform upsampling This module further improves segmentation by explicitly exploiting information interaction between densely and sparsely resampled areas at multiple scales. Both modules can seamlessly integrate with various architectures, extending from convolutional neural networks to transformers. Feature visualization and analysis demonstrate that our method effectively alleviates aliasing while successfully retaining details after demodulation. As a result, the proposed approach considerably enhances existing state-of-the-art segmentation models (e.g., Mask2Former-Swin-T +1.5 mIoU, InternImage-T +1.4 mIoU on ADE20 K). Furthermore, ARS also enhances the performance of powerful Deformable Convolution (+0.8 mIoU on Cityscapes) by maintaining relative positional order during non-uniform sampling. Finally, we validate the broad applicability and effectiveness of SFM by extending it to image classification, adversarial robustness, instance segmentation, and panoptic segmentation tasks.
Ying Fu 0001, Lin Gu 0003, Dezhi Zheng, Jifeng Dai
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Investigating Synthetic-to-Real Transfer Robustness for Stereo Matching and Optical Flow Estimation
abstract
With advancements in robust stereo matching and optical flow estimation networks, models pre-trained on synthetic data demonstrate strong robustness to unseen domains. However, their robustness can be seriously degraded when fine-tuning them in real-world scenarios. This paper investigates fine-tuning stereo matching and optical flow estimation networks without compromising their robustness to unseen domains. Specifically, we divide the pixels into consistent and inconsistent regions by comparing Ground Truth (GT) with Pseudo Label (PL) and demonstrate that the imbalance learning of consistent and inconsistent regions in GT causes robustness degradation. Based on our analysis, we propose the DKT framework, which utilizes PL to balance the learning of different regions in GT. The core idea is to utilize an exponential moving average (EMA) teacher to measure what the student network has learned and dynamically adjust the learning regions. We further propose the DKT++ framework, which improves target-domain performances and network robustness by applying slow-fast update teachers to generate more accurate PL, introducing the unlabeled data and synthetic data. We integrate our frameworks with state-of-the-art networks and evaluate their effectiveness on several real-world datasets. Extensive experiments show that our method effectively preserves the robustness of stereo matching and optical flow networks during fine-tuning.
Jiahe Li 0007, Lei Huang 0015, Haonan Luo 0002, Xiaohan Yu 0001, Lin Gu 0003, Xiao Bai 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 A New Benchmark: Clinical Uncertainty and Severity Aware Labeled Chest X-Ray Images With Multi-Relationship Graph Learning
abstract
Chest radiography, commonly known as CXR, is frequently utilized in clinical settings to detect cardiopulmonary conditions. However, even seasoned radiologists might offer different evaluations regarding the seriousness and uncertainty associated with observed abnormalities. Previous research has attempted to utilize clinical notes to extract abnormal labels for training deep-learning models in CXR image diagnosis. However, these methods often neglected the varying degrees of severity and uncertainty linked to different labels. In our study, we initially assembled a comprehensive new dataset of CXR images based on clinical textual data, which incorporated radiologists' assessments of uncertainty and severity. Using this dataset, we introduced a multi-relationship graph learning framework that leverages spatial and semantic relationships while addressing expert uncertainty through a dedicated loss function. Our research showcases a notable enhancement in CXR image diagnosis and the interpretability of the diagnostic model, surpassing existing state-of-the-art methodologies. The dataset address of disease severity and uncertainty we extracted is: https://physionet.org/content/cad-chest/1.0/.
Mengliang Zhang, Xinyue Hu 0002, Lin Gu 0003, Kazuma Kobayashi, Tatsuya Harada, Ronald M. Summers, Yingying Zhu 0004
IEEE Trans. Medical Imaging3
2024 Aleth-NeRF: Illumination Adaptive NeRF with Concealing Field Assumption
abstract
The standard Neural Radiance Fields (NeRF) paradigm employs a viewer-centered methodology, entangling the aspects of illumination and material reflectance into emission solely from 3D points. This simplified rendering approach presents challenges in accurately modeling images captured under adverse lighting conditions, such as low light or over-exposure. Motivated by the ancient Greek emission theory that posits visual perception as a result of rays emanating from the eyes, we slightly refine the conventional NeRF framework to train NeRF under challenging light conditions and generate normal-light condition novel views unsupervisedly. We introduce the concept of a ``Concealing Field," which assigns transmittance values to the surrounding air to account for illumination effects. In dark scenarios, we assume that object emissions maintain a standard lighting level but are attenuated as they traverse the air during the rendering process. Concealing Field thus compel NeRF to learn reasonable density and colour estimations for objects even in dimly lit situations. Similarly, the Concealing Field can mitigate over-exposed emissions during rendering stage. Furthermore, we present a comprehensive multi-view dataset captured under challenging illumination conditions for evaluation. Our code and proposed dataset are available at https://github.com/cuiziteng/Aleth-NeRF.
Ziteng Cui, Lin Gu 0003, Xiao Sun 0001, Xianzheng Ma, Yu Qiao 0001, Tatsuya Harada
AAAI2
2024 Discovering an Image-Adaptive Coordinate System for Photography Processing
Ziteng Cui, Lin Gu 0003, Tatsuya Harada
BMVC2
2024 DNGaussian: Optimizing Sparse-View 3D Gaussian Radiance Fields with Global-Local Depth Normalization
abstract
Radiance fields have demonstrated impressive performance in synthesizing novel views from sparse input views, yet prevailing methods suffer from high training costs and slow inference speed. This paper introduces DNGaussian, a depth-regularized framework based on 3D Gaussian radiance fields, offering real-time and high-quality few-shot novel view synthesis at low costs. Our motivation stems from the highly efficient representation and surprising quality of the recent 3D Gaussian Splatting, despite it will encounter a geometry degradation when input views decrease. In the Gaussian radiance fields, we find this degradation in scene geometry primarily lined to the positioning of Gaussian primitives and can be mitigated by depth constraint. Consequently, we propose a Hard and Soft Depth Regularization to restore accurate scene geometry under coarse monocular depth supervision while maintaining a fine-grained color appearance. To further refine detailed geometry reshaping, we introduce Global-Local Depth Normalization, enhancing the focus on small local depth changes. Extensive experiments on LLFF, DTU, and Blender datasets demonstrate that DNGaussian outperforms state-of-the-art methods, achieving comparable or better results with significantly reduced memory cost, a 25 × reduction in training time, and over 3000 × faster rendering speed. Code is available at: https://github.com/Fictionarry/DNGaussian.
Jiahe Li 0007, Xiao Bai 0001, Xin Ning 0001, Jun Zhou 0001, Lin Gu 0003
CVPR7
2024 Frequency-Adaptive Dilated Convolution for Semantic Segmentation
abstract
Dilated convolution, which expands the receptive field by inserting gaps between its consecutive elements, is widely employed in computer vision. In this study, we propose three strategies to improve individual phases of dilated convolution from the perspective of spectrum analysis. Departing from the conventional practice of fixing a global dilation rate as a hyperparameter, we introduce Frequency-Adaptive Dilated Convolution (FADC), which dynamically adjusts dilation rates spatially based on local frequency components. Subsequently, we design two plug-in modules to directly enhance effective bandwidth and receptive field size. The Adaptive Kernel (AdaKern) module decomposes convolution weights into low-frequency and high-frequency components, dynamically adjusting the ratio between these components on a per-channel basis. By increasing the high-frequency part of convolution weights, AdaKern captures more high-frequency components, thereby improving effective bandwidth. The Frequency Selection (FreqSelect) module optimally balances high- and low-frequency components in feature representations through spatially variant reweighting. It suppresses high frequencies in the background to encourage FADC to learn a larger dilation, thereby increasing the receptive field for an expanded scope. Extensive experiments on segmentation and object detection consistently validate the efficacy of our approach. The code is made publicly available at https://github.com/ying-fu/FADC.
Lin Gu 0003, Dezhi Zheng, Ying Fu 0001
CVPR2
2024 Robust Synthetic-to-Real Transfer for Stereo Matching
abstract
With advancements in domain generalized stereo matching networks, models pre-trained on synthetic data demonstrate strong robustness to unseen domains. However, few studies have investigated the robustness after fine-tuning them in real-world scenarios, during which the domain generalization ability can be seriously degraded. In this paper, we explore fine-tuning stereo matching networks without compromising their robustness to unseen domains. Our motivation stems from comparing Ground Truth (GT) versus Pseudo Label (PL) for fine-tuning: GT degrades, but PL preserves the domain generalization ability. Empirically, we find the difference between GT and PL implies valuable information that can regularize networks during fine-tuning. We also propose a framework to utilize this difference for fine-tuning, consisting of a frozen Teacher, an exponential moving average (EMA) Teacher, and a Student network. The core idea is to utilize the EMA Teacher to measure what the Student has learned and dynamically improve GT and PL for fine-tuning. We integrate our framework with state-of-the-art networks and evaluate its effectiveness on several real-world datasets. Extensive experiments show that our method effectively preserves the domain generalization ability during fine-tuning. Code is available at: https://github.com/jiaw-z/DKT-Stereo.
Jiahe Li 0007, Lei Huang 0015, Xiaohan Yu 0001, Lin Gu 0003, Xiao Bai 0001
CVPR5
2024 Object-Aware NIR-to-Visible Translation
Yunyi Gao, Lin Gu 0003, Qiankun Liu 0001, Ying Fu 0001
ECCV (23)2
2024 TalkingGaussian: Structure-Persistent 3D Talking Head Synthesis via Gaussian Splatting
Jiahe Li 0007, Xiao Bai 0001, Xin Ning 0001, Jun Zhou 0001, Lin Gu 0003
ECCV (10)7
2024 CoR-GS: Sparse-View 3D Gaussian Splatting via Co-regularization
Jiahe Li 0007, Xiaohan Yu 0001, Lei Huang 0015, Lin Gu 0003, Xiao Bai 0001
ECCV (1)5
2024 When Semantic Segmentation Meets Frequency Aliasing
abstract
Despite recent advancements in semantic segmentation, where and what pixels are hard to segment remains largely unexplored. Existing research only separates an image into easy and hard regions and empirically observes the latter are associated with object boundaries. In this paper, we conduct a comprehensive analysis of hard pixel errors, categorizing them into three types: false responses, merging mistakes, and displacements. Our findings reveal a quantitative association between hard pixels and aliasing, which is distortion caused by the overlapping of frequency components in the Fourier domain during downsampling. To identify the frequencies responsible for aliasing, we propose using the equivalent sampling rate to calculate the Nyquist frequency, which marks the threshold for aliasing. Then, we introduce the aliasing score as a metric to quantify the extent of aliasing. While positively correlated with the proposed aliasing score, three types of hard pixels exhibit different patterns. Here, we propose two novel de-aliasing filter (DAF) and frequency mixing (FreqMix) modules to alleviate aliasing degradation by accurately removing or adjusting frequencies higher than the Nyquist frequency. The DAF precisely removes the frequencies responsible for aliasing before downsampling, while the FreqMix dynamically selects high-frequency components within the encoder block. Experimental results demonstrate consistent improvements in semantic segmentation and low-light instance segmentation tasks. The code is at: \url{https://github.com/Linwei-Chen/Seg-Aliasing}.
Lin Gu 0003, Ying Fu 0001
ICLR2
2024 TinyLUT: Tiny Look-Up Table for Efficient Image Restoration at the Edge
abstract
Look-up tables(LUTs)-based methods have recently shown enormous potential in image restoration tasks, which are capable of significantly accelerating the inference. However, the size of LUT exhibits exponential growth with the convolution kernel size, creating a storage bottleneck for its broader application on edge devices. Here, we address the storage explosion challenge to promote the capacity of mapping the complex CNN models by LUT. We introduce an innovative separable mapping strategy to achieve over $7\times$ storage reduction, transforming the storage from exponential dependence on kernel size to a linear relationship. Moreover, we design a dynamic discretization mechanism to decompose the activation and compress the quantization scale that further shrinks the LUT storage by $4.48\times$. As a result, the storage requirement of our proposed TinyLUT is around 4.1\% of MuLUT-SDY-X2 and amenable to on-chip cache, yielding competitive accuracy with over $5\times$ lower inference latency on Raspberry 4B than FSRCNN. Our proposed TinyLUT enables superior inference speed on edge devices with new state-of-the-art accuracy on both of image super-resolution and denoising, showcasing the potential of applying this method to various image restoration tasks at the edge. The codes are available at: https://github.com/Jonas-KD/TinyLUT.
Huanan Li, Juntao Guan, Lai Rui, Sijun Ma, Lin Gu 0003, Noperson
NeurIPS5
2024 SRMAE: Masked Image Modeling for Scale-Invariant Deep Representations
Lin Gu 0003, Feng Lu 0005
PRCV (2)2
2024 Can physician judgment enhance model trustworthiness? A case study on predicting pathological lymph nodes in rectal cancer
abstract
Explainability is key to enhancing the trustworthiness of artificial intelligence in medicine. However, there exists a significant gap between physicians' expectations for model explainability and the actual behavior of these models. This gap arises from the absence of a consensus on a physician-centered evaluation framework, which is needed to quantitatively assess the practical benefits that effective explainability should offer practitioners. Here, we hypothesize that superior attention maps, as a mechanism of model explanation, should align with the information that physicians focus on, potentially reducing prediction uncertainty and increasing model reliability. We employed a multimodal transformer to predict lymph node metastasis of rectal cancer using clinical data and magnetic resonance imaging. We explored how well attention maps, visualized through a state-of-the-art technique, can achieve agreement with physician understanding. Subsequently, we compared two distinct approaches for estimating uncertainty: a standalone estimation using only the variance of prediction probability, and a human-in-the-loop estimation that considers both the variance of prediction probability and the quantified agreement. Our findings revealed no significant advantage of the human-in-the-loop approach over the standalone one. In conclusion, this case study did not confirm the anticipated benefit of the explanation in enhancing model reliability. Superficial explanations could do more harm than good by misleading physicians into relying on uncertain predictions, suggesting that the current state of attention mechanisms should not be overestimated in the context of model explainability.
Kazuma Kobayashi, Yasuyuki Takamizawa, Mototaka Miyake, Sono Ito, Lin Gu 0003, Tatsuya Nakatsuka, Yu Akagi, Tatsuya Harada, Yukihide Kanemitsu, Ryuji Hamamoto
Artif. Intell. Medicine5
2024 Exploring the Usage of Pre-trained Features for Stereo Matching
Lei Huang 0015, Xiao Bai 0001, Lin Gu 0003, Edwin R. Hancock
Int. J. Comput. Vis.5
2024 Interpretable medical image Visual Question Answering via multi-modal relationship graph learning
Xinyue Hu 0002, Lin Gu 0003, Kazuma Kobayashi, Mengliang Zhang, Tatsuya Harada, Ronald M. Summers, Yingying Zhu 0003
Medical Image Anal.2
2024 Sketch-based semantic retrieval of medical images
abstract
The volume of medical images stored in hospitals is rapidly increasing; however, the utilization of these accumulated medical images remains limited. Existing content-based medical image retrieval (CBMIR) systems typically require example images, leading to practical limitations, such as the lack of customizable, fine-grained image retrieval, the inability to search without example images, and difficulty in retrieving rare cases. In this paper, we introduce a sketch-based medical image retrieval (SBMIR) system that enables users to find images of interest without the need for example images. The key concept is feature decomposition of medical images, which allows the entire feature of a medical image to be decomposed into and reconstructed from normal and abnormal features. Building on this concept, our SBMIR system provides an easy-to-use two-step graphical user interface: users first select a template image to specify a normal feature and then draw a semantic sketch of the disease on the template image to represent an abnormal feature. The system integrates both types of input to construct a query vector and retrieves reference images. For evaluation, ten healthcare professionals participated in a user test using two datasets. Consequently, our SBMIR system enabled users to overcome previous challenges, including image retrieval based on fine-grained image characteristics, image retrieval without example images, and image retrieval for rare cases. Our SBMIR system provides on-demand, customizable medical image retrieval, thereby expanding the utility of medical image databases.
Kazuma Kobayashi, Lin Gu 0003, Ryuichiro Hataya, Takaaki Mizuno, Mototaka Miyake, Hirokazu Watanabe, Masamichi Takahashi, Yasuyuki Takamizawa, Yukihiro Yoshida, Nobuji Kouno, Amina Bolatkan, Yusuke Kurose, Tatsuya Harada, Ryuji Hamamoto
Medical Image Anal.2
2024 Rethinking masked image modelling for medical image representation
abstract
Masked Image Modelling (MIM), a form of self-supervised learning, has garnered significant success in computer vision by improving image representations using unannotated data. Traditional MIMs typically employ a strategy of random sampling across the image. However, this random masking technique may not be ideally suited for medical imaging, which possesses distinct characteristics divergent from natural images. In medical imaging, particularly in pathology, disease-related features are often exceedingly sparse and localized, while the remaining regions appear normal and undifferentiated. Additionally, medical images frequently accompany reports, directly pinpointing pathological changes' location. Inspired by this, we propose Masked medical Image Modelling (MedIM), a novel approach, to our knowledge, the first research that employs radiological reports to guide the masking and restore the informative areas of images, encouraging the network to explore the stronger semantic representations from medical images. We introduce two mutual comprehensive masking strategies, knowledge-driven masking (KDM), and sentence-driven masking (SDM). KDM uses Medical Subject Headings (MeSH) words unique to radiology reports to identify symptom clues mapped to MeSH words (e.g., cardiac, edema, vascular, pulmonary) and guide the mask generation. Recognizing that radiological reports often comprise several sentences detailing varied findings, SDM integrates sentence-level information to identify key regions for masking. MedIM reconstructs images informed by this masking from the KDM and SDM modules, promoting a comprehensive and enriched medical image representation. Our extensive experiments on seven downstream tasks covering multi-label/class image classification, pneumothorax segmentation, and medical image-report analysis, demonstrate that MedIM with report-guided masking achieves competitive performance. Our method substantially outperforms ImageNet pre-training, MIM-based pre-training, and medical image-report pre-training counterparts. Codes are available at https://github.com/YtongXie/MedIM.
Yutong Xie 0001, Lin Gu 0003, Tatsuya Harada, Yong Xia 0001, Qi Wu 0001
Medical Image Anal.2
2024 Learning From Human Attention for Attribute-Assisted Visual Recognition
abstract
With prior knowledge of seen objects, humans have a remarkable ability to recognize novel objects using shared and distinct local attributes. This is significant for the challenging tasks of zero-shot learning (ZSL) and fine-grained visual classification (FGVC), where the discriminative attributes of objects have played an important role. Inspired by human visual attention, neural networks have widely exploited the attention mechanism to learn the locally discriminative attributes for challenging tasks. Though greatly promoted the development of these fields, existing works mainly focus on learning the region embeddings of different attribute features and neglect the importance of discriminative attribute localization. It is also unclear whether the learned attention truly matches the real human attention. To tackle this problem, this paper proposes to employ real human gaze data for visual recognition networks to learn from human attention. Specifically, we design a unified Attribute Attention Network (A$^{2}$Net) that learns from human attention for both ZSL and FGVC tasks. The overall model consists of an attribute attention branch and a baseline classification network. On top of the image feature maps provided by the baseline classification network, the attribute attention branch employs attribute prototypes to produce attribute attention maps and attribute features. The attribute attention maps are converted to gaze-like attentions to be aligned with real human gaze attention. To guarantee the effectiveness of attribute feature learning, we further align the extracted attribute features with attribute-defined class embeddings. To facilitate learning from human gaze attention for the visual recognition problems, we design a bird classification game to collect real human gaze data using the CUB dataset via an eye-tracker device. Experiments on ZSL and FGVC tasks without/with real human gaze data validate the benefits and accuracy of our proposed model. This work supports the promising benefits of collecting human gaze datasets and automatic gaze estimation algorithms learning from human attention for high-level computer vision tasks.
Xiao Bai 0001, Pengcheng Zhang 0003, Xiaohan Yu 0001, Edwin R. Hancock, Jun Zhou 0001, Lin Gu 0003
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 Frequency-Aware Feature Fusion for Dense Image Prediction
abstract
Dense image prediction tasks demand features with strong category information and precise spatial boundary details at high resolution. To achieve this, modern hierarchical models often utilize feature fusion, directly adding upsampled coarse features from deep layers and high-resolution features from lower levels. In this paper, we observe rapid variations in fused feature values within objects, resulting in intra-category inconsistency due to disturbed high-frequency features. Additionally, blurred boundaries in fused features lack accurate high frequency, leading to boundary displacement. Building upon these observations, we propose Frequency-Aware Feature Fusion (FreqFusion), integrating an Adaptive Low-Pass Filter (ALPF) generator, an offset generator, and an Adaptive High-Pass Filter (AHPF) generator. The ALPF generator predicts spatially-variant low-pass filters to attenuate high-frequency components within objects, reducing intra-class inconsistency during upsampling. The offset generator refines large inconsistent features and thin boundaries by replacing inconsistent features with more consistent ones through resampling, while the AHPF generator enhances high-frequency detailed boundary information lost during downsampling. Comprehensive visualization and quantitative analysis demonstrate that FreqFusion effectively improves feature consistency and sharpens object boundaries. Extensive experiments across various dense prediction tasks confirm its effectiveness.
Ying Fu 0001, Lin Gu 0003, Chenggang Yan 0001, Tatsuya Harada, Gao Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 People Taking Photos That Faces Never Share: Privacy Protection and Fairness Enhancement from Camera to User
abstract
The soaring number of personal mobile devices and public cameras poses a threat to fundamental human rights and ethical principles. For example, the stolen of private information such as face image by malicious third parties will lead to catastrophic consequences. By manipulating appearance of face in the image, most of existing protection algorithms are effective but irreversible. Here, we propose a practical and systematic solution to invertiblely protect face information in the full-process pipeline from camera to final users. Specifically, We design a novel lightweight Flow-based Face Encryption Method (FFEM) on the local embedded system privately connected to the camera, minimizing the risk of eavesdropping during data transmission. FFEM uses a flow-based face encoder to encode each face to a Gaussian distribution and encrypts the encoded face feature by random rotating the Gaussian distribution with the rotation matrix is as the password. While encrypted latent-variable face images are sent to users through public but less reliable channels, password will be protected through more secure channels through technologies such as asymmetric encryption, blockchain, or other sophisticated security schemes. User could select to decode an image with fake faces from the encrypted image on the public channel. Only trusted users are able to recover the original face using the encrypted matrix transmitted in secure channel. More interestingly, by tuning Gaussian ball in latent space, we could control the fairness of the replaced face on attributes such as gender and race. Extensive experiments demonstrate that our solution could protect privacy and enhance fairness with minimal effect on high-level downstream task.
Lin Gu 0003, Sissi Xiaoxiao Wu, Tatsuya Harada, Yingying Zhu 0004
AAAI2
2023 Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait Synthesis
abstract
This paper presents ER-NeRF, a novel conditional Neural Radiance Fields (NeRF) based architecture for talking portrait synthesis that can concurrently achieve fast convergence, real-time rendering, and state-of-the-art performance with small model size. Our idea is to explicitly exploit the unequal contribution of spatial regions to guide talking portrait modeling. Specifically, to improve the accuracy of dynamic head reconstruction, a compact and expressive NeRF-based Tri-Plane Hash Representation is introduced by pruning empty spatial regions with three planar hash encoders. For speech audio, we propose a Region Attention Module to generate region-aware condition feature via an attention mechanism. Different from existing methods that utilize an MLP-based encoder to learn the cross-modal relation implicitly, the attention mechanism builds an explicit connection between audio features and spatial regions to capture the priors of local motions. Moreover, a direct and fast Adaptive Pose Encoding is introduced to optimize the head-torso separation problem by mapping the complex transformation of the head pose into spatial coordinates. Extensive experiments demonstrate that our method renders better high-fidelity and audio-lips synchronized talking portrait videos, with realistic details and high efficiency compared to previous methods. Code is available at https://github.com/Fictionarry/ER-NeRF.
Jiahe Li 0007, Xiao Bai 0001, Jun Zhou 0001, Lin Gu 0003
ICCV5
2023 Name Your Colour For the Task: Artificially Discover Colour Naming via Colour Quantisation Transformer
abstract
The long-standing theory that a colour-naming system evolves under dual pressure of efficient communication and perceptual mechanism is supported by more and more linguistic studies, including analysing four decades of diachronic data from the Nafaanra language. This inspires us to explore whether machine learning could evolve and discover a similar colour-naming system via optimising the communication efficiency represented by high-level recognition performance. Here, we propose a novel colour quantisation transformer, CQFormer, that quantises colour space while maintaining the accuracy of machine recognition on the quantised images. Given an RGB image, Annotation Branch maps it into an index map before generating the quantised image with a colour palette; meanwhile the Palette Branch utilises a key-point detection way to find proper colours in the palette among the whole colour space. By interacting with colour annotation, CQFormer is able to balance both the machine vision accuracy and colour perceptual structure such as distinct and stable colour distribution for discovered colour system. Very interestingly, we even observe the consistent evolution pattern between our artificial colour system and basic colour terms across human languages. Besides, our colour quantisation method also offers an efficient quantisation method that effectively compresses the image storage while maintaining high performance in high-level recognition tasks such as classification and detection. Extensive experiments demonstrate the superior performance of our method with extremely low bit-rate colours, showing potential to integrate into quantisation network to quantities from image to network activation. The source code is available at https://github.com/ryeocthiv/CQFormer
Shenghan Su, Lin Gu 0003, Zenghui Zhang, Tatsuya Harada
ICCV2
2023 3D Segmenter: 3D Transformer based Semantic Segmentation via 2D Panoramic Distillation
Zhennan Wu, Yang Li 0193, Yifei Huang 0002, Lin Gu 0003, Tatsuya Harada, Hiroyuki Sato 0002
ICLR4
2023 Expert Knowledge-Aware Image Difference Graph Representation Learning for Difference-Aware Medical Visual Question Answering
abstract
To contribute to automating the medical vision-language model, we propose a novel Chest-Xray Different Visual Question Answering (VQA) task. Given a pair of main and reference images, this task attempts to answer several questions on both diseases and, more importantly, the differences between them. This is consistent with the radiologist's diagnosis practice that compares the current image with the reference before concluding the report. We collect a new dataset, namely MIMIC-Diff-VQA, including 700,703 QA pairs from 164,324 pairs of main and reference images. Compared to existing medical VQA datasets, our questions are tailored to the Assessment-Diagnosis-Intervention-Evaluation treatment procedure used by clinical professionals. Meanwhile, we also propose a novel expert knowledge-aware graph representation learning model to address this task. The proposed baseline model leverages expert knowledge such as anatomical structure prior, semantic, and spatial knowledge to construct a multi-relationship graph, representing the image differences between two images for the image difference VQA task. The dataset and code can be found at https://github.com/Holipori/MIMIC-Diff-VQA. We believe this work would further push forward the medical vision language model.
Xinyue Hu 0002, Lin Gu 0003, Qiyuan An, Mengliang Zhang, Kazuma Kobayashi, Tatsuya Harada, Ronald M. Summers, Yingying Zhu 0003
KDD2
2023 Towards AI-Driven Radiology Education: A Self-supervised Segmentation-Based Framework for High-Precision Medical Image Editing
Kazuma Kobayashi, Lin Gu 0003, Ryuichiro Hataya, Mototaka Miyake, Yasuyuki Takamizawa, Sono Ito, Hirokazu Watanabe, Yukihiro Yoshida, Hiroki Yoshimura, Tatsuya Harada, Ryuji Hamamoto
MICCAI (2)2
2023 MedIM: Boost Medical Image Representation via Radiology Report-Guided Masking
Yutong Xie 0001, Lin Gu 0003, Tatsuya Harada, Yong Xia 0001, Qi Wu 0001
MICCAI (1)2
2023 Information bottleneck and selective noise supervision for zero-shot learning
Lei Zhou 0008, Yang Liu 0357, Pengcheng Zhang 0003, Xiao Bai 0001, Lin Gu 0003, Jun Zhou 0001, Yazhou Yao, Tatsuya Harada, Edwin R. Hancock
Mach. Learn.5
2023 Correlated and individual feature learning with contrast-enhanced MR for malignancy characterization of hepatocellular carcinoma
abstract
Malignancy characterization of hepatocellular carcinoma (HCC) is of great importance in patient management and prognosis prediction. In this study, we propose an end-to-end correlated and individual feature learning framework to characterize the malignancy of HCC from Contrast-enhanced MR. From the phases of pre-contrast, arterial and portal venous, our framework simultaneously and explicitly learns both the shareable and phase-specific features that are discriminative to malignancy grades. We evaluate our method on the Contrast enhanced MR of 112 consecutive patients with 117 histologically proven HCCs. Experimental results demonstrate that arterial phase yields better results than portal vein and pre-contrast phase. Furthermore, phase specific components show better discriminant ability than the shareable components. Finally, combining the extracted shareable and individual features components has yielded significantly better performance than traditional feature fusion methods. We also conduct t-SNE analysis and feature scoring analysis to qualitatively assess the effectiveness of the proposed method for malignancy characterization.
Yunling Li, Shangxuan Li, Hanqiu Ju, Tatsuya Harada, Honglai Zhang, Ting Duan, Guangyi Wang, Lin Gu 0003, Wu Zhou 0002
Pattern Recognit.9
2023 DnRCNN: Deep Recurrent Convolutional Neural Network for HSI Destriping
abstract
In spite of achieving promising results in hyperspectral image (HSI) restoration, deep-learning-based methodologies still face the problem of spectral or spatial information loss due to neglecting the inner correlation of HSI. To address this issue, we propose an innovative deep recurrent convolution neural network (DnRCNN) model for HSI destriping. To the best of our knowledge, this is the first study on HSI destriping from the perspective of inner band and interband correlation explorations with the recurrent convolution neural network. In the novel DnRCNN, a selective recurrent memory unit (SRMU) is designed to respectively extract the correlative features involved in spectral and spatial domains. Moreover, an innovative recurrent fusion (RF) strategy incorporated with group concatenation is further proposed to remove strip noise and preserve scene details using the complementary features from SRMU. Experimental results on extensive HSI datasets validated that the proposed method achieves a new state-of-the-art (SOTA) HSI destriping performance.
Juntao Guan, Huanan Li, Yintang Yang, Lin Gu 0003
IEEE Trans. Neural Networks Learn. Syst.5
2023 Multiresolution Discriminative Mixup Network for Fine-Grained Visual Categorization
abstract
Fine-grained visual categorization (FGVC) is a challenging task because there are many hard examples existing between fine-grained classes which differ subtly in particular local regions. To address this issue, many methods have recourse to high-resolution source images and others adopt effective regularization like "mixup" or "between class learning." Despite their promising achievements, mixup tends to cause the manifold intrusion problem which would result in under-fitting and degradation of the model performance and high-resolution input inevitably leads to high computational costs. In view of this, we present a multiresolution discriminative mixup network (MRDMN). Different from standard mixup, the proposed discriminative mixup strategy mixes discriminative regions linearly instead of entire images to avoid manifold intrusion, which makes it learn the local detail features more effectively and contributes to more precise categorization. Furthermore, an innovative resolution-based distillation strategy is designed to transfer the multiresolution detail feature representations to a low-resolution network, which speeds up the testing and boosts the categorization accuracy simultaneously. Extensive experiments demonstrate that our proposed MRDMN remarkably outperforms most competitive approaches with less computation time on the CUB-200-2011, Stanford-Cars, Stanford-Dogs, Food-101, and iNaturalist 2017 datasets. The codes are in https://github.com/aztc/MRDMN.
Kunran Xu, Lin Gu 0003, Yishi Li
IEEE Trans. Neural Networks Learn. Syst.3
2022 Towards an Effective Orthogonal Dictionary Convolution Strategy
abstract
Orthogonality regularization has proven effective in improving the precision, convergence speed and the training stability of CNNs. Here, we propose a novel Orthogonal Dictionary Convolution Strategy (ODCS) on CNNs to improve orthogonality effect by optimizing the network architecture and changing the regularized object. Specifically, we remove the nonlinear layer in typical convolution block “Conv(BN) + Nonlinear + Pointwise Conv(BN)”, and only impose orthogonal regularization on the front Conv. The structure, “Conv(BN) + Pointwise Conv(BN)”, is then equivalent to a pair of dictionary and encoding, defined in sparse dictionary learning. Thanks to the exact and efficient representation of signal with dictionaries in low-dimensional projections, our strategy could reduce the superfluous information in dictionary Conv kernels. Meanwhile, the proposed strategy relieves the too strict orthogonality regularization in training, which makes hyper-parameters tuning of model to be more flexible. In addition, our ODCS can modify the state-of-the-art models easily without any extra consumption in inference phase. We evaluate it on a variety of CNNs in small-scale (CIFAR), large-scale (ImageNet) and fine-grained (CUB-200-2011) image classification tasks, respectively. The experimental results show that our method achieve a stable and superior improvement.
Yishi Li, Kunran Xu, Lin Gu 0003
AAAI4
2022 EtinyNet: Extremely Tiny Network for TinyML
abstract
There are many AI applications in high-income countries because their implementation depends on expensive GPU cards (~2000$) and reliable power supply (~200W). To deploy AI in resource-poor settings on cheaper (~20$) and low-power devices (
Kunran Xu, Yishi Li, Lin Gu 0003
AAAI5
2022 You Only Need 90K Parameters to Adapt Light: a Light Weight Transformer for Image Enhancement and Exposure Correction
Ziteng Cui, Kunchang Li 0002, Lin Gu 0003, Shenghan Su, Peng Gao 0007, Zhengkai Jiang 0001, Yu Qiao 0001, Tatsuya Harada
BMVC3
2022 Revisiting Domain Generalized Stereo Matching Networks from a Feature Consistency Perspective
abstract
Despite recent stereo matching networks achieving impressive performance given sufficient training data, they suffer from domain shifts and generalize poorly to unseen domains. We argue that maintaining feature consistency between matching pixels is a vital factor for promoting the generalization capability of stereo matching networks, which has not been adequately considered. Here we address this issue by proposing a simple pixel-wise contrastive learning across the viewpoints. The stereo contrastive feature loss function explicitly constrains the consistency between learned features of matching pixel pairs which are observations of the same 3D points. A stereo selective whitening loss is further introduced to better preserve the stereo feature consistency across domains, which decorrelates stereo features from stereo viewpoint-specific style information. Counter-intuitively, the generalization of feature consistency between two viewpoints in the same scene translates to the generalization of stereo matching performance to unseen domains. Our method is generic in nature as it can be easily embedded into existing stereo networks and does not require access to the samples in the target domain. When trained on synthetic data and generalized to four real-world testing sets, our method achieves superior performance over several state-of-the-art networks. The code is available online11https://github.com/jiaw-z/FCStereo.
Xiang Wang 0014, Xiao Bai 0001, Chen Wang 0026, Lei Huang 0015, Lin Gu 0003, Jun Zhou 0001, Tatsuya Harada, Edwin R. Hancock
CVPR7
2022 Exploring Resolution and Degradation Clues as Self-supervised Signal for Low Quality Object Detection
Ziteng Cui, Yingying Zhu 0004, Lin Gu 0003, Guo-Jun Qi, Renrui Zhang, Zenghui Zhang, Tatsuya Harada
ECCV (9)3
2022 Where to Focus: Investigating Hierarchical Attention Relationship for Fine-Grained Visual Classification
Yang Liu 0357, Lei Zhou 0008, Pengcheng Zhang 0003, Xiao Bai 0001, Lin Gu 0003, Xiaohan Yu 0001, Jun Zhou 0001, Edwin R. Hancock
ECCV (24)5
2022 Surgical Skill Assessment via Video Semantic Aggregation
Zhenqiang Li 0002, Lin Gu 0003, Weimin Wang 0007, Ryosuke Nakamura, Yoichi Sato 0001
MICCAI (8)2
2022 A Geometry-Constrained Deformable Attention Network for Aortic Segmentation
Weiyuan Lin, Lin Gu 0003, Zhifan Gao
MICCAI (5)3
2022 A multilayer dynamic perturbation analysis method for predicting ligand-protein interactions
abstract
BACKGROUND: Ligand-protein interactions play a key role in defining protein function, and detecting natural ligands for a given protein is thus a very important bioengineering task. In particular, with the rapid development of AI-based structure prediction algorithms, batch structural models with high reliability and accuracy can be obtained at low cost, giving rise to the urgent requirement for the prediction of natural ligands based on protein structures. In recent years, although several structure-based methods have been developed to predict ligand-binding pockets and ligand-binding sites, accurate and rapid methods are still lacking, especially for the prediction of ligand-binding regions and the spatial extension of ligands in the pockets. RESULTS: In this paper, we proposed a multilayer dynamics perturbation analysis (MDPA) method for predicting ligand-binding regions based solely on protein structure, which is an extended version of our previously developed fast dynamic perturbation analysis (FDPA) method. In MDPA/FDPA, ligand binding tends to occur in regions that cause large changes in protein conformational dynamics. MDPA, examined using a standard validation dataset of ligand-protein complexes, yielded an averaged ligand-binding site prediction Matthews coefficient of 0.40, with a prediction precision of at least 50% for 71% of the cases. In particular, for 80% of the cases, the predicted ligand-binding region overlaps the natural ligand by at least 50%. The method was also compared with other state-of-the-art structure-based methods. CONCLUSIONS: MDPA is a structure-based method to detect ligand-binding regions on protein surface. Our calculations suggested that a range of spaces inside the protein pockets has subtle interactions with the protein, which can significantly impact on the overall dynamics of the protein. This work provides a valuable tool as a starting point upon which further docking and analysis methods can be used for natural ligand detection in protein functional annotation. The source code of MDPA method is freely available at: https://github.com/mingdengming/mdpa .
Lin Gu 0003, Dengming Ming
BMC Bioinform.1
2022 Memory-Efficient Deformable Convolution Based Joint Denoising and Demosaicing for UHD Images
abstract
This paper introduces deformable convolution in deep learning based joint denoising and demosaicing (JDD), which yields more adaptable representation and larger receptive fields in features extraction for a superior restoration performance. However, the deformable convolution generally leads to considerable computational load and irregular memory access bottleneck, limiting its extensive deployment on edge devices. To address this issue, we develop grouping strategy and assign independent offsets to each kernel group to reduce the computation latency while keeping the accuracy. Motivated by the exploration for aggregate distribution characteristics of deformable offsets, we present the offset sharing methodology to simplify the memory access complexity of deformable convolution. As for hardware acceleration, we specially design a novel deformable matrix multiplication workflow incorporated with a deformable memory mapping unit to boost the computational throughput. The verification experiments on FPGA demonstrate that the proposed deformable convolution based JDD can restore 4K Ultra High Definition (UHD) images at 70FPS and yields significant promotion in visual effect and objective quality assessment.
Juntao Guan, Yangang Li, Huanan Li, Lichen Feng, Yintang Yang, Lin Gu 0003
IEEE Trans. Circuits Syst. Video Technol.8
2022 Explainable Diabetic Retinopathy Detection and Retinal Image Generation
abstract
Though deep learning has shown successful performance in classifying the label and severity stage of certain diseases, most of them give few explanations on how to make predictions. Inspired by Koch's Postulates, the foundation in evidence-based medicine (EBM) to identify the pathogen, we propose to exploit the interpretability of deep learning application in medical diagnosis. By isolating neuron activation patterns from a diabetic retinopathy (DR) detector and visualizing them, we can determine the symptoms that the DR detector identifies as evidence to make prediction. To be specific, we first define novel pathological descriptors using activated neurons of the DR detector to encode both spatial and appearance information of lesions. Then, to visualize the symptom encoded in the descriptor, we propose Patho-GAN, a new network to synthesize medically plausible retinal images. By manipulating these descriptors, we could even arbitrarily control the position, quantity, and categories of generated lesions. We also show that our synthesized images carry the symptoms directly related to diabetic retinopathy diagnosis. Our generated images are both qualitatively and quantitatively superior to the ones by previous methods. Besides, compared to existing methods that take hours to generate an image, our second level speed endows the potential to be an effective solution for data augmentation.
Yuhao Niu, Lin Gu 0003, Yitian Zhao, Feng Lu 0005
IEEE J. Biomed. Health Informatics2
2022 Beyond Triplet Loss: Person Re-Identification With Fine-Grained Difference-Aware Pairwise Loss
abstract
Person Re-IDentification (ReID) aims at re-identifying persons from different viewpoints across multiple cameras. Capturing the fine-grained appearance differences is often the key to accurate person ReID, because many identities can be differentiated only when looking into these fine-grained differences. However, most state-of-the-art person ReID approaches, typically driven by a triplet loss, fail to effectively learn the fine-grained features as they are focused more on differentiating large appearance differences. To address this issue, we introduce a novel pairwise loss function that enables ReID models to learn the fine-grained features by adaptively enforcing an exponential penalization on the images of small differences and a bounded penalization on the images of large differences. The proposed loss is generic and can be used as a plugin to replace the triplet loss to significantly enhance different types of state-of-the-art approaches. Experimental results on four benchmark datasets show that the proposed loss substantially outperforms a number of popular loss functions by large margins; and it also enables significantly improved data efficiency.
Guansong Pang, Xiao Bai 0001, Changhong Liu, Xin Ning 0001, Lin Gu 0003, Jun Zhou 0001
IEEE Trans. Multim.6
2021 Leveraging Human Selective Attention for Medical Image Analysis with Limited Training Data
Yifei Huang 0002, Lijin Yang, Lin Gu 0003, Yingying Zhu 0004, Hirofumi Seo, Qiuming Meng, Tatsuya Harada, Yoichi Sato 0001
BMVC4
2021 Goal-Oriented Gaze Estimation for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) aims to recognize novel classes by transferring semantic knowledge from seen classes to unseen classes. Since semantic knowledge is built on attributes shared between different classes, which are highly local, strong prior for localization of object attribute is beneficial for visual-semantic embedding. Interestingly, when recognizing unseen images, human would also automatically gaze at regions with certain semantic clue. Therefore, we introduce a novel goal-oriented gaze estimation module (GEM) to improve the discriminative attribute localization based on the class-level attributes for ZSL. We aim to predict the actual human gaze location to get the visual attention regions for recognizing a novel object guided by attribute description. Specifically, the task-dependent attention is learned with the goal-oriented GEM, and the global image features are simultaneously optimized with the regression of local attribute features. Experiments on three ZSL benchmarks, i.e., CUB, SUN and AWA2, show the superiority or competitiveness of our proposed method against the state-of-the-art ZSL methods. The ablation analysis on real gaze data CUB-VWSW also validates the benefits and accuracy of our gaze estimation module. This work implies the promising benefits of collecting human gaze dataset and automatic gaze estimation algorithms on high-level computer vision tasks. The code is available at https://github.com/osierboy/GEM-ZSL.
Yang Liu 0357, Lei Zhou 0008, Xiao Bai 0001, Yifei Huang 0002, Lin Gu 0003, Jun Zhou 0001, Tatsuya Harada
CVPR5
2021 Multitask AET with Orthogonal Tangent Regularity for Dark Object Detection
abstract
Dark environment becomes a challenge for computer vision algorithms owing to insufficient photons and undesirable noise. To enhance object detection in a dark environment, we propose a novel multitask auto encoding transformation (MAET) model which is able to explore the intrinsic pattern behind illumination translation. In a self-supervision manner, the MAET learns the intrinsic visual structure by encoding and decoding the realistic illumination-degrading transformation considering the physical noise model and image signal processing (ISP). Based on this representation, we achieve the object detection task by decoding the bounding box coordinates and classes. To avoid the over-entanglement of two tasks, our MAET disentangles the object and degrading features by imposing an orthogonal tangent regularity. This forms a parametric manifold along which multitask predictions can be geometrically formulated by maximizing the orthogonality between the tangents along the outputs of respective tasks. Our framework can be implemented based on the mainstream object detection architecture and directly trained end-to-end using normal target detection datasets, such as VOC and COCO. We have achieved the state-of-the-art performance using synthetic and real-world datasets. Codes will be released at https://github.com/cuiziteng/MAET.
Ziteng Cui, Guo-Jun Qi, Lin Gu 0003, Shaodi You, Zenghui Zhang, Tatsuya Harada
ICCV3
2021 Relation-Aware Reasoning with Graph Convolutional Network
Lei Zhou 0008, Yang Liu 0357, Xiao Bai 0001, Xiang Wang 0014, Chen Wang 0026, Liang Zhang 0044, Lin Gu 0003
ICIG (1)7
2021 Understanding adversarial attacks on deep learning based medical image analysis systems
Xingjun Ma, Yuhao Niu, Lin Gu 0003, Yisen Wang 0001, Yitian Zhao, James Bailey 0001, Feng Lu 0005
Pattern Recognit.3
2020 Feature Normalized Knowledge Distillation for Image Classification
Kunran Xu, Lai Rui, Yishi Li, Lin Gu 0003
ECCV (25)4
2020 Motion Feedback Design for Video Frame Interpolation
abstract
This paper introduces a feedback-based approach to interpolate video frames involving small and fast-moving objects. Unlike the existing feedforward-based methods that estimate optical flow and synthesize in-between frames sequentially, we introduce a motion-oriented component that adds a feedback block to the existing multi-scale autoencoder pipeline, which feedbacks information of small objects shared between architectures of two different scales. We show that feeding this additional information enables more robust detection of optical flow caused by small objects in fast motion. Using experiments on various datasets, we show that the feedback mechanism allows our method to achieve state-of-the-art results, both qualitatively and quantitatively.
Mengshun Hu, Jing Xiao 0004, Lin Gu 0003, Shin'ichi Satoh 0001
ICASSP4
2020 Fixed pattern noise reduction for infrared images based on cascade residual attention CNN
Juntao Guan, Ai Xiong, Zesheng Liu, Lin Gu 0003
Neurocomputing5
2020 Hyperspectral Anomaly Detection via Tensor- Based Endmember Extraction and Low-Rank Decomposition
abstract
Due to the limited resolution of hyperspectral sensors, anomalous targets expressed with subpixels are often mixed with nonhomogeneous backgrounds. This fact makes anomalies difficult to distinguish from the surrounding background. From this perspective, a novel hyperspectral anomaly detection (AD) algorithm based on endmember extraction and low-rank representation (LRR) is proposed. For the characteristics of pixels in hyperspectral images (HSIs), the proposed algorithm employs an endmember extraction technology to yield an abundance matrix for AD, thereby gathering more feature information compared with the direct use of a raw image. In addition, a dictionary construction strategy based on Tucker decomposition, and the${k}$-means++ clustering method is proposed to make the dictionary more stable and discriminative. An LRR method based on the dictionary is applied to obtain a sparse residual matrix. Finally, anomalies can be determined by the response of the residual matrix. Experiments on three hyperspectral data sets validate the performance of the proposed algorithm.
Shangzhen Song, Huixin Zhou, Lin Gu 0003, Yixin Yang 0002, Yiyi Yang
IEEE Geosci. Remote. Sens. Lett.3
2020 Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark Study
abstract
Existing enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions.
Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin
IEEE Trans. Image Process.40
2019 Pathological Evidence Exploration in Deep Retinal Image Diagnosis
abstract
Though deep learning has shown successful performance in classifying the label and severity stage of certain disease, most of them give few evidence on how to make prediction. Here, we propose to exploit the interpretability of deep learning application in medical diagnosis. Inspired by Koch’s Postulates, a well-known strategy in medical research to identify the property of pathogen, we define a pathological descriptor that can be extracted from the activated neurons of a diabetic retinopathy detector. To visualize the symptom and feature encoded in this descriptor, we propose a GAN based method to synthesize pathological retinal image given the descriptor and a binary vessel segmentation. Besides, with this descriptor, we can arbitrarily manipulate the position and quantity of lesions. As verified by a panel of 5 licensed ophthalmologists, our synthesized images carry the symptoms that are directly related to diabetic retinopathy diagnosis. The panel survey also shows that our generated images is both qualitatively and quantitatively superior to existing methods.
Yuhao Niu, Lin Gu 0003, Feng Lu 0005, Feifan Lv, Zongji Wang, Imari Sato, Zijian Zhang 0004, Yangyan Xiao, Xunzhang Dai
AAAI2
2019 Discrepancy Steered Conditional Adversarial Network For Deep Feature Based Malignancy Characterization of Hepatocellular Carcinoma
abstract
Preoperative knowledge of the malignancy of hepatocellular carcinoma (HCC) based on medical images plays a significant role in deciding therapy strategies and patient management in clinical prac-tice. Deep learning with Convolutional Neural Network (CNN) has shown high diagnostic performance for lesion characterization with medical images. However, it is very challenging to train robust deep learning system for lesion characterization, especially for HCC, because there are often limited samples in clinical practice. In this work, we propose an efficient end-to-end framework that flexibly combines Conditional Adversarial network (CAN) and CNN to characterize the malignancy of HCC. Specifically, we introduce a similarity discriminative network to make the CAN efficiently generate more discrepant samples and devise hybrid loss functions to embed the proposed similarity discriminative network to the end-to-end framework with CAN and CNN. Experimental results of 115 clinical HCCs with pathologically confirmed malignancy demonstrate that the proposed end-to-end framework with similarity discriminative network can significantly improve the performance of deep feature based malignancy characterization of HCC and remarkably reduce the risk of overfitting with limited samples for the deep learning model in clinical practice.
Hanqiu Ju, Guangyi Wang, Shaoyang Men, Honglai Zhang, Lin Gu 0003, Wu Zhou 0002
ICIP5
2019 Unsupervised Ensemble Strategy for Retinal Vessel Segmentation
Bo Liu 0112, Lin Gu 0003, Feng Lu 0005
MICCAI (1)2
2018 A Data-Driven Approach for Direct and Global Component Separation from a Single Image
Shijie Nie, Lin Gu 0003, Art Subpa-Asa, Ilyes Kacher, Ko Nishino, Imari Sato
ACCV (6)2
2018 Deeply Learned Filter Response Functions for Hyperspectral Reconstruction
abstract
Hyperspectral reconstruction from RGB imaging has recently achieved significant progress via sparse coding and deep learning. However, a largely ignored fact is that existing RGB cameras are tuned to mimic human trichromatic perception, thus their spectral responses are not necessarily optimal for hyperspectral reconstruction. In this paper, rather than use RGB spectral responses, we simultaneously learn optimized camera spectral response functions (to be implemented in hardware) and a mapping for spectral reconstruction by using an end-to-end network. Our core idea is that since camera spectral filters act in effect like the convolution layer, their response functions could be optimized by training standard neural networks. We propose two types of designed filters: a three-chip setup without spatial mosaicing and a single-chip setup with a Bayer-style 2x2 filter array. Numerical simulations verify the advantages of deeply learned spectral responses compared to existing RGB cameras. More interestingly, by considering physical restrictions in the design process, we are able to realize the deeply learned spectral response functions by using modern film filter production technologies, and thus construct data-inspired multispectral cameras for snapshot hyperspectral imaging.
Shijie Nie, Lin Gu 0003, Yinqiang Zheng, Antony Lam, Nobutaka Ono, Imari Sato
CVPR2
2017 From RGB to Spectrum for Natural Scenes via Manifold-Based Mapping
abstract
Spectral analysis of natural scenes can provide much more detailed information about the scene than an ordinary RGB camera. The richer information provided by hyperspectral images has been beneficial to numerous applications, such as understanding natural environmental changes and classifying plants and soils in agriculture based on their spectral properties. In this paper, we present an efficient manifold learning based method for accurately reconstructing a hyperspectral image from a single RGB image captured by a commercial camera with known spectral response. By applying a nonlinear dimensionality reduction technique to a large set of natural spectra, we show that the spectra of natural scenes lie on an intrinsically low dimensional manifold. This allows us to map an RGB vector to its corresponding hyperspectral vector accurately via our proposed novel manifold-based reconstruction pipeline. Experiments using both synthesized RGB images using hyperspectral datasets and real world data demonstrate our method outperforms the state-of-the-art.
Yan Jia 0005, Yinqiang Zheng, Lin Gu 0003, Art Subpa-Asa, Antony Lam, Yoichi Sato 0001, Imari Sato
ICCV3
2017 Semi-supervised Learning for Biomedical Image Segmentation via Forest Oriented Super Pixels(Voxels)
Lin Gu 0003, Yinqiang Zheng, Ryoma Bise, Imari Sato, Nobuaki Imanishi, Sadakazu Aiso
MICCAI (1)1
2017 Segment 2D and 3D Filaments by Learning Structured and Contextual Features
abstract
We focus on the challenging problem of filamentary structure segmentation in both 2D and 3D images, including retinal vessels and neurons, among others. Despite the increasing amount of efforts in learning based methods to tackle this problem, there still lack proper data-driven feature construction mechanisms to sufficiently encode contextual labelling information, which might hinder the segmentation performance. This observation prompts us to propose a data-driven approach to learn structured and contextual features in this paper. The structured features aim to integrate local spatial label patterns into the feature space, thus endowing the follow-up tree classifiers capability to grouping training examples with similar structure into the same leaf node when splitting the feature space, and further yielding contextual features to capture more of the global contextual information. Empirical evaluations demonstrate that our approach outperforms state-of-the-arts on well-regarded testbeds over a variety of applications. Our code is also made publicly available in support of the open-source research activities.
Lin Gu 0003, Xiaowei Zhang 0002, He Zhao 0002, Huiqi Li, Li Cheng 0001
IEEE Trans. Medical Imaging1
2016 A quadratic optimisation approach for shading and specularity recovery from a single image
abstract
In this paper we present a method to recover the shading and specularities in the scene from a single image. The method presented here is based on the dichromatic model and enforces a local smoothness assumption over the object surfaces in the scene. This naturally leads to a setting where the estimate of the shading at a particular pixel can be expressed in terms of its neighbours up to a pair of Gaussian kernels accounting for the irradiance similarity between pixels and their spatial proximity on the image plane. This yields a quadratic cost function for both, the specular coefficient and the shading factor of the dicromatic model which can be solved using gradient descent. We show results for both, specular highlight recovery and shading estimation and compare them against a number of alternatives.
Lin Gu 0003, Antonio Robles-Kelly
ICIP1
2016 Classification from a Riemannian graph embedding viewpoint
abstract
In this paper, we employ graph embeddings for classification tasks. To do this, we explore the relationship between kernel matrices, spaces of inner products and statistical inference by viewing the embedding vectors for the nodes in the graph as a field on a Riemannian manifold. This leads to a setting where the inference process may be cast as a Maximum a Posteriori (MAP) estimation over a Gibbs field whereby the graph Laplacian can be related to a Gram matrix of scalar products. This not only allows for a better understanding of graph spectral techniques, but also provides a means for classifying nodes in the graph without the need to compute the embedding explicitly by using a Mercer kernel. We illustrate how the developments presented here can be used for purposes of classification, where we use the graph Laplacian as a kernel matrix. We present classification results on synthetic data and four UCI datasets. We also apply our method to real-world image labelling and compare our results to those yielded by alternatives elsewhere in the literature.
Antonio Robles-Kelly, Lin Gu 0003
IJCNN2
2015 Learning to Boost Filamentary Structure Segmentation
abstract
The challenging problem of filamentary structure segmentation has a broad range of applications in biological and medical fields. A critical yet challenging issue remains on how to detect and restore the small filamentary fragments from backgrounds: The small fragments are of diverse shapes and appearances, meanwhile the backgrounds could be cluttered and ambiguous. Focusing on this issue, this paper proposes an iterative two-step learning-based approach to boost the performance based on a base segmenter arbitrarily chosen from a number of existing segmenters: We start with an initial partial segmentation where the filamentary structure obtained is of high confidence based on this existing segmenter. We also define a scanning horizon as epsilon balls centred around the partial segmentation result. Step one of our approach centers on a data-driven latent classification tree model to detect the filamentary fragments. This model is learned via a training process, where a large number of distinct local figure/background separation scenarios are established and geometrically organized into a tree structure. Step two spatially restores the isolated fragments back to the current partial segmentation, which is accomplished by means of completion fields and matting. Both steps are then alternated with the growth of partial segmentation result, until the input image space is entirely explored. Our approach is rather generic and can be easily augmented to a wide range of existing supervised/unsupervised segmenters to produce an improved result. This has been empirically verified on specific filamentary structure segmentation tasks: retinal blood vessel segmentation as well as neuronal segmentations, where noticeable improvement has been shown over the original state-of-the-arts.
Lin Gu 0003, Li Cheng 0002
ICCV1
2014 Shadow modelling based upon Rayleigh scattering and Mie theory
Lin Gu 0003, Antonio Robles-Kelly
Pattern Recognit. Lett.1
2014 Segmentation and Estimation of Spatially Varying Illumination
abstract
In this paper, we present an unsupervised method for segmenting the illuminant regions and estimating the illumination power spectrum from a single image of a scene lit by multiple light sources. Here, illuminant region segmentation is cast as a probabilistic clustering problem in the image spectral radiance space. We formulate the problem in an optimization setting, which aims to maximize the likelihood of the image radiance with respect to a mixture model while enforcing a spatial smoothness constraint on the illuminant spectrum. We initialize the sample pixel set under each illuminant via a projection of the image radiance spectra onto a low-dimensional subspace spanned by a randomly chosen subset of spectra. Subsequently, we optimize the objective function in a coordinate-ascent manner by updating the weights of the mixture components, sample pixel set under each illuminant, and illuminant posterior probabilities. We then estimate the illuminant power spectrum per pixel making use of these posterior probabilities. We compare our method with a number of alternatives for the tasks of illumination region segmentation, illumination color estimation, and color correction. Our experiments show the effectiveness of our method as applied to one hyperspectral and three trichromatic image data sets.
Lin Gu 0003, Cong Phuoc Huynh, Antonio Robles-Kelly
IEEE Trans. Image Process.1
2013 Efficient Estimation of Reflectance Parameters From Imaging Spectroscopy
abstract
In this paper, we address the problem of efficiently recovering reflectance parameters from a single multispectral or hyperspectral image. To do so, we propose a shapelet based estimator that employs shapelets to recover the shading in the image. The optimization setting presented is based upon a three-step process. The first of these concerns the recovery of the surface reflectance and the specular coefficients through a constrained optimization approach. Second, we update the illuminant power spectrum using a simple least-squares formulation. Third, the shading is computed directly once the updated illuminant power spectrum is obtained. This yields a computationally efficient method that achieves speed-ups of nearly an order of magnitude over its closest alternative without compromising performance. We provide results on illuminant power spectrum computation, shading recovery, skin recognition and replacement of the scene illuminant, and object reflectance in real-world images.
Lin Gu 0003, Antonio Robles-Kelly, Jun Zhou 0001
IEEE Trans. Image Process.1
2012 Shadow detection via Rayleigh scattering and Mie theory
Lin Gu 0003, Antonio Robles-Kelly
ICPR1
2011 Material-specific user colour profiles from imaging spectroscopy data
abstract
In this paper, we present a method which permits the creation of user colour preferences for object materials and lights in the scene making use of imaging spectroscopy data. To do this, we build upon the heterogeneous nature of the scene by imposing consistency over object materials so as to allow for small compositional variations across objects in the image. Once the consistency has been imposed, we aim at maximising the quality of the images under consideration based upon user input. This provides the flexibility necessary to utilise user profiles for the automatic processing of real world imagery while avoiding undesirable effects encountered when colour images are produced. We provide results on real-world imagery and illustrate how the method can be used to produce material-specific colours based upon user input.
Lin Gu 0003, Cong Phuoc Huynh, Antonio Robles-Kelly, Jun Zhou 0001
ICCV1