EDBT 2026 Demo / reviewers in the wild / expert
Yu Guo 0008
dblp:53/382-8
· DBLP profile ↗
22ranked-venue papers
8as first author
22since 2021 · last 2026
0000-0002-0642-7684ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QMSANet: A quaternion multi-scale attention network for robust color image denoising
Qi Xie 0002, Yu Guo 0008, Boying Wu, Deyu Meng, Jean-Michel Morel, Qiyu Jin, Michael Kwok-Po Ng |
Neural Networks | 3 |
| 2026 | Quaternion adaptive approximation normalization graph guided implicit low rank for robust matrix completion
Yu Guo 0008, Tieyong Zeng, Qiyu Jin, Michael Kwok-Po Ng |
Pattern Recognit. | 1 |
| 2025 | Task-Oriented Communications for Visual Navigation with Edge-Aerial Collaboration in Low Altitude EconomyabstractTo support the development of the Low Altitude Economy (LAE), it is essential to achieve precise localization of unmanned aerial vehicles (UAVs) in urban areas where global positioning system (GPS) signals are unavailable. Vision-based methods offer a viable alternative but face severe bandwidth, memory and processing constraints on lightweight UAVs. Inspired by mammalian spatial cognition, we propose a task-oriented communication framework, where UAVs equipped with multi-camera systems extract compact multi-view features and offload localization tasks to edge servers. We introduce the Orthogonally-constrained Variational Information Bottleneck encoder (O-VIB), which incorporates automatic relevance determination (ARD) to prune non-informative features while enforcing orthogonality to minimize redundancy. This enables efficient and accurate localization with minimal transmission cost. Extensive evaluation on a dedicated LAE UAV dataset shows that O-VIB achieves high-precision localization under stringent bandwidth budgets. Code and dataset will be made publicly available: github.com/fangzr/TOC-Edge-Aerial. Zhengru Fang, Jingjing Wang 0001, Senkang Hu, Yu Guo 0008, Yiqin Deng, Yuguang Fang |
GLOBECOM | 5 |
| 2025 | Instruct2See: Learning to Remove Any Obstructions Across DistributionsabstractImages are often obstructed by various obstacles due to capture limitations, hindering the observation of objects of interest. Most existing methods address occlusions from specific elements like fences or raindrops, but are constrained by the wide range of real-world obstructions, making comprehensive data collection impractical. To overcome these challenges, we propose Instruct2See, a novel zero-shot framework capable of handling both seen and unseen obstacles. The core idea of our approach is to unify obstruction removal by treating it as a soft-hard mask restoration problem, where any obstruction can be represented using multi-modal prompts, such as visual semantics and textual instructions, processed through a cross-attention unit to enhance contextual understanding and improve mode control. Additionally, a tunable mask adapter allows for dynamic soft masking, enabling real-time adjustment of inaccurate masks. Extensive experiments on both in-distribution and out-of-distribution obstacles show that Instruct2See consistently achieves strong performance and generalization in obstruction removal, regardless of whether the obstacles were present during the training phase. Code and dataset are available at https://jhscut.github.io/Instruct2See. Junhang Li, Yu Guo 0008, Chuhua Xian, Shengfeng He |
ICML | 2 |
| 2025 | Neptune-X: Active X-to-Maritime Generation for Universal Maritime Object DetectionabstractMaritime object detection is essential for navigation safety, surveillance, and autonomous operations, yet constrained by two key challenges: the scarcity of annotated maritime data and poor generalization across various maritime attributes (e.g., object category, viewpoint, location, and imaging environment). To address these challenges, we propose Neptune-X, a data-centric generative-selection framework that enhances training effectiveness by leveraging synthetic data generation with task-aware sample selection. From the generation perspective, we develop X-to-Maritime, a multi-modality-conditioned generative model that synthesizes diverse and realistic maritime scenes. A key component is the Bidirectional Object-Water Attention module, which captures boundary interactions between objects and their aquatic surroundings to improve visual fidelity. To further improve downstream tasking performance, we propose Attribute-correlated Active Sampling, which dynamically selects synthetic samples based on their task relevance. To support robust benchmarking, we construct the Maritime Generation Dataset, the first dataset tailored for generative maritime learning, encompassing a wide range of semantic conditions. Extensive experiments demonstrate that our approach sets a new benchmark in maritime scene synthesis, significantly improving detection accuracy, particularly in challenging and previously underrepresented settings. The code is available at https://github.com/gy65896/Neptune-X. Yu Guo 0008, Shengfeng He, Yuxu Lu, Haonan An 0001, Yihang Tao, Huilin Zhu, Jingxian Liu, Yuguang Fang |
NeurIPS | 1 |
| 2025 | Quaternion deep matrix factorization and non-local Laplacian regularization for matrix completion
Yu Guo 0008, Qiyu Jin, Tieyong Zeng, Michael Kwok-Po Ng |
Knowl. Based Syst. | 1 |
| 2025 | Quaternion Nuclear Norm Minus Frobenius Norm Minimization for color image reconstruction
Yu Guo 0008, Tieyong Zeng, Qiyu Jin, Michael Kwok-Po Ng |
Pattern Recognit. | 1 |
| 2024 | OneRestore: A Universal Restoration Framework for Composite Degradation
Yu Guo 0008, Yuan Gao 0015, Yuxu Lu, Huilin Zhu, Ryan Wen Liu, Shengfeng He |
ECCV (19) | 1 |
| 2024 | Zero-Shot Object Counting with Good Exemplars
Huilin Zhu, Jingling Yuan, Zhengwei Yang 0001, Yu Guo 0008, Zheng Wang 0007, Xian Zhong, Shengfeng He |
ECCV (5) | 4 |
| 2024 | AoSRNet: All-in-One Scene Recovery Networks via multi-knowledge integration
Yuxu Lu, Dong Yang 0003, Yuan Gao 0015, Ryan Wen Liu, Jun Liu 0012, Yu Guo 0008 |
Knowl. Based Syst. | 6 |
| 2024 | Kernel correlation-dissimilarity for Multiple Kernel k-Means clustering
Rina Su, Yu Guo 0008, Caiying Wu, Qiyu Jin, Tieyong Zeng |
Pattern Recognit. | 2 |
| 2024 | Deep Inertia $L_{p}$ Half-Quadratic Splitting Unrolling Network for Sparse View CT ReconstructionabstractSparse view computed tomography (CT) reconstruction poses a challenging ill-posed inverse problem, necessitating effective regularization techniques. In this letter, we employ$L_{p}$-norm ($0< p< 1$) regularization to induce sparsity and introduce inertial steps, leading to the development of the inertial$L_{p}$-norm half-quadratic splitting algorithm. We rigorously prove the convergence of this algorithm. Furthermore, we leverage deep learning to initialize the conjugate gradient method, resulting in a deep unrolling network with theoretical guarantees. Our extensive numerical experiments demonstrate that our proposed algorithm surpasses existing methods, particularly excelling in fewer scanned views and complex noise conditions. Yu Guo 0008, Caiying Wu, Qiyu Jin, Tieyong Zeng |
IEEE Signal Process. Lett. | 1 |
| 2024 | Real-Time Multi-Scene Visibility Enhancement for Promoting Navigational Safety of Vessels Under Complex Weather ConditionsabstractThe visible-light camera, which is capable of environment perception and navigation assistance, has emerged as an essential imaging sensor for marine surface vessels in intelligent waterborne transportation systems (IWTS). However, the visual imaging quality inevitably suffers from several kinds of degradations (e.g., limited visibility, low contrast, color distortion, etc.) under complex weather conditions (e.g., haze, rain, and low-lightness). The degraded visual information will accordingly result in inaccurate environment perception and delayed operations for navigational risk. To promote the navigational safety of vessels, many computational methods have been presented to perform visual quality enhancement under poor weather conditions. However, most of these methods are essentially specific-purpose implementation strategies, only available for one specific weather type. To overcome this limitation, we propose to develop a general-purpose multi-scene visibility enhancement method, i.e., edge reparameterization- and attention-guided neural network (ERANet), to adaptively restore the degraded images captured under different weather conditions. In particular, our ERANet simultaneously exploits the channel attention, spatial attention, and reparameterization technology to enhance the visual quality while maintaining low computational cost. Extensive experiments conducted on standard and IWTS-related datasets have demonstrated that our ERANet could outperform several representative visibility enhancement methods in terms of both imaging quality and computational efficiency. The superior performance of IWTS-related object detection and scene segmentation could also be steadily obtained after ERANet-based visibility enhancement under complex weather conditions. Ryan Wen Liu, Yuxu Lu, Yuan Gao 0015, Yu Guo 0008, Wenqi Ren, Fenghua Zhu, Fei-Yue Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Double Domain Guided Real-Time Low-Light Image Enhancement for Ultra-High-Definition Transportation SurveillanceabstractReal-time transportation surveillance is an essential part of the intelligent transportation system (ITS). However, images captured under low-light conditions often suffer poor visibility with types of degradation, such as noise interference and vague edge features, etc. With the development of imaging devices, the quality of the visual surveillance data is continually increasing, like 2K and 4K, which have more strict requirements on the efficiency of image processing. To satisfy the requirements on both enhancement quality and computational speed, this paper proposes a double domain guided real-time low-light image enhancement network (DDNet) for ultra-high-definition (UHD) transportation surveillance. Specifically, we design an encoder-decoder structure as the main architecture of the learning network. In particular, the enhancement processing is divided into two subtasks (i.e., color enhancement and gradient enhancement) via the proposed coarse enhancement module (CEM) and LoG-based gradient enhancement module (GEM), which are embedded in the encoder-decoder structure. It enables the network to enhance the color and edge features simultaneously. Through the decomposition and reconstruction on both color and gradient domains, our DDNet can restore the detailed feature information concealed by the darkness with better visual quality and efficiency. The evaluation experiments on standard and transportation-related datasets demonstrate that our DDNet provides superior enhancement quality and efficiency compared with state-of-the-art methods. Besides, the object detection and scene segmentation experiments indicate the practical benefits for higher-level image analysis under low-light environments in ITS. The source code is available at https://github.com/QuJX/DDNet. Jingxiang Qu, Ryan Wen Liu, Yuan Gao 0015, Yu Guo 0008, Fenghua Zhu, Fei-Yue Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Multi-Task Learning-Enabled Automatic Vessel Draft Reading for Intelligent Maritime SurveillanceabstractThe accurate and efficient vessel draft reading (VDR) is an important component of intelligent maritime surveillance, which could be exploited to assist in judging whether the vessel is normally loaded or overloaded. The computer vision technique with an excellent price-to-performance ratio has become a popular medium to estimate vessel draft depth. However, the traditional estimation methods easily suffer from several limitations, such as sensitivity to low-quality images, high computational cost, etc. In this work, we propose a multi-task learning-enabled computational method (termed MTL-VDR) for generating highly reliable VDR. In particular, our MTL-VDR mainly consists of four components, i.e., draft mark detection, draft scale recognition, vessel/water segmentation, and final draft depth estimation. We first construct a benchmark dataset related to draft mark detection and employ a powerful and efficient convolutional neural network to accurately perform the detection task. The multi-task learning method is then proposed for simultaneous draft scale recognition and vessel/water segmentation. To obtain more robust VDR under complex conditions (e.g., damaged and stained scales, etc.), the accurate draft scales are generated by an automatic correction method, which is presented based on the spatial distribution rules of draft scales. Finally, an adaptive computational method is exploited to yield an accurate and robust draft depth. Extensive experiments have been implemented on the realistic dataset to compare our MTL-VDR with state-of-the-art methods. The results have demonstrated its superior performance in terms of accuracy, robustness, and efficiency. The computational speed exceeds 40 FPS, which satisfies the requirements of real-time maritime surveillance to guarantee vessel traffic safety. Jingxiang Qu, Ryan Wen Liu, Chenjie Zhao, Yu Guo 0008, Sendren Sheng-Dong Xu, Fenghua Zhu |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Deep Network-Enabled Haze Visibility Enhancement for Visual IoT-Driven Intelligent Transportation SystemsabstractThe Internet of Things (IoT) has recently emerged as a revolutionary communication paradigm where a large number of objects and devices are closely interconnected to enable smart industrial environments. The tremendous growth of visual sensors can significantly promote the traffic situational awareness, traffic safety management, and intelligent vehicle navigation in intelligent transportation systems (ITSs). However, due to the absorption and scattering of light by the turbid medium in atmosphere, the visual IoT inevitably suffers from imaging quality degradation, e.g., contrast reduction, color distortion, etc. This negative impact can not only reduce the imaging quality, but also bring challenges for the deployment of several high-level vision tasks (e.g., object detection, tracking, recognition, etc.) in the ITS. To improve imaging quality under the hazy environment, we propose a deep network-enabled three-stage dehazing network (termed TSDNet) for promoting the visual IoT-driven ITS. In particular, the proposed TSDNet mainly contains three parts, i.e., multiscale attention module for estimating the hazy distribution in the RGB image domain, two-branch extraction module for learning the hazy features, and multifeature fusion module for integrating all characteristic information and reconstructing the haze-free image. Numerous experiments have been implemented on synthetic and real-world imaging scenarios. Dehazing results illustrated that our TSDNet remarkably outperformed several state-of-the-art methods in terms of both qualitative and quantitative evaluations. The high-accuracy object detection results have also demonstrated the superior dehazing performance of the TSDNet under hazy atmosphere conditions. The source code is available athttps://github.com/gy65896/TSDNet. Ryan Wen Liu, Yu Guo 0008, Yuxu Lu, Kwok Tai Chui, Brij B. Gupta |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | GradDT: Gradient-Guided Despeckling Transformer for Industrial Imaging SensorsabstractThe speckle noise is a granular disturbance that often brings negative side effects on the detection and recognition of targets of interest in industrial imaging sensors. From the statistical point of view, this type of noise can be modeled as a multiplicative formula. The nonlinear multiplicative property makes despeckling more intractable with respect to noise reduction and details preservation. To blindly remove the undesirable speckle noise, we combine the gradient model and machine learning technology for despeckling. In particular, we first introduce the logarithmic transformation to transform the multiplicative speckle noise into an additive version. A gradient-guided despeckling transformer (termed GradDT) is then proposed to blindly reduce the additive noise in the transformed noisy images. To be specific, the proposed method mainly includes two modules, i.e., the spatial feature extraction module (SFEM) and the efficient transformer module (ETM). The SFEM can extract the spatial feature of speckle noise and the gradient maps corresponding to the noise-free image. The ETM module can calculate the spatial domain's cross-channel cross-covariance and produce global attention maps to reconstruct the sharp image. The proposed GradDT thus can effectively distinguish the speckle noise and vital image features (e.g., edge and texture) to balance the degree of noise suppression and details preservation. Extensive experiments have been implemented on both synthetic and realistic degraded images. Compared with several state-of-the-art speckle noise reduction methods, our GradDT could generate superior imaging performance in terms of both quantitative evaluation and visual quality. Yuxu Lu, Yu Guo 0008, Ryan Wen Liu, Kwok Tai Chui, Brij B. Gupta |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | Asynchronous Trajectory Matching-Based Multimodal Maritime Data Fusion for Vessel Traffic Surveillance in Inland WaterwaysabstractThe automatic identification system (AIS) and video cameras have been widely exploited for vessel traffic surveillance in inland waterways. The AIS data could provide vessel identity and dynamic information on vessel position and movements. In contrast, the video data could describe the visual appearances of moving vessels without knowing the information on identity, position, movements, etc. To further improve vessel traffic surveillance, it becomes necessary to fuse the AIS and video data to simultaneously capture the visual features, identity, and dynamic information for the vessels of interest. However, the performance of AIS and video data fusion is susceptible to issues such as data spatial difference, message asynchronous transmission, visual object occlusion, etc. In this work, we propose a deep learning-based simple online and real-time vessel data fusion method (termed DeepSORVF). We first extract the AIS-and video-based vessel trajectories, and then propose an asynchronous trajectory matching method to fuse the AIS-based vessel information with the corresponding visual targets. In addition, by combining the AIS-and video-based movement features, we also present a prior knowledge-driven anti-occlusion method to yield accurate and robust vessel tracking results under occlusion conditions. To validate the efficacy of our DeepSORVF, we have also constructed a new benchmark dataset (termed FVessel) for vessel detection, tracking, and data fusion. It consists of many videos and the corresponding AIS data collected in various weather conditions and locations. The experimental results have demonstrated that our method is capable of guaranteeing high-reliable data fusion and anti-occlusion vessel tracking. The DeepSORVF code and FVessel dataset are publicly available at https://github.com/gy65896/DeepSORVF and https://github.com/gy65896/FVessel, respectively. Yu Guo 0008, Ryan Wen Liu, Jingxiang Qu, Yuxu Lu, Fenghua Zhu |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Rep-Enhancer: Re-parameterizing Neural Network for Real-time Low-light Enhancement in Visual Maritime SurveillanceabstractVision-based maritime surveillance has become an essential part of the vessel traffic services system. The images collected in low-light maritime conditions often suffer from poor visibility. These images may significantly degenerate the performance of high-level visual tasks and increase the uncertainty in maritime surveillance. To address this problem, we propose a lightweight neural network (Rep-Enhancer) for low-light image enhancement. Specifically, we first design a re-parameterizable multi-branch edge extraction module, i.e., spatial domain-oriented convolution block (SDCB). Furthermore, skip connections and spatial attention operations are employed to strengthen the features. By exploiting these well-strengthened edge features, we can enhance the low-light images effectively with the encoder-decoder structure. The experimental results have shown that Rep-Enhancer can enhance the low-light image qualifiedly while maintaining great inference efficiency. Xijing Li, Yuxu Lu, Yu Guo 0008, Jingxiang Qu, Ryan Wen Liu |
EUC | 3 |
| 2022 | Gaussian Patch Mixture Model Guided Low-Rank Covariance Matrix Minimization for Image DenoisingabstractImage denoising is one of the most important tasks in image processing. In this paper, we study image denoising methods by using similar patches which have low-rank covariance matrices to recover an underlying image which is corrupted by additive Gaussian noise. In order to enhance global patch-matching results, we make use of a Gaussian mixture model with an auxiliary image to determine different groups of patches. The auxiliary image is an output of BM3D. The noisy version of covariance matrix is formed by each group of patches from the given noisy image. Its low-rank version can be estimated by using covariance matrix nuclear norm minimization, and the resulting denoised image can be obtained. Experimental results are reported to show that the proposed method outperforms the state-of-the-art denoising methods, including testing deep learning methods, in the peak signal-to-noise ratio, structural similarity values, and visual quality. Yu Guo 0008, Qiyu Jin, Michael Kwok-Po Ng |
SIAM J. Imaging Sci. | 2 |
| 2022 | MTRBNet: Multi-Branch Topology Residual Block-Based Network for Low-Light EnhancementabstractThe learning-based low-light image enhancement methods have remarkable performance due to the robust feature learning and mapping capabilities. This paper proposes a multi-branch topology residual block (MTRB)-based network (MTRBNet), which can alleviate training difficulties and more efficiently use the parameters between neurons. Compared with the previous residual block, the proposed MTRB increases the width of the network and simultaneously transmits information along with the depth and width directions, which can effectively select network nodes to promote the network learning capacity. Meanwhile, the feature information of neighbor nodes is transferred to each other, thereby maximizing the information flow of the convolution unit. The proposed information connection and feedback mechanism can improve the network’s ability to capture the global and local features. We analyze the pros and cons of two multi-feature fusion strategies (i.e., addition and concatenation) and three normalization methods on the quantitative results. In addition, we embed our MTRB into traditional Encoder-Decoder structure to improve the image enhancement results under different low-light imaging conditions. Experiments on the LOL image dataset have demonstrated that our MTRBNet achieves superior performance compared with several state-of-the-art methods. Yuxu Lu, Yu Guo 0008, Ryan Wen Liu, Wenqi Ren |
IEEE Signal Process. Lett. | 2 |
| 2021 | Fast, Nonlocal and Neural: A Lightweight High Quality Solution to Image DenoisingabstractWith the widespread application of convolutional neural networks (CNNs), the traditional model based denoising algorithms are now outperformed. However, CNNs face two problems. First, they are computationally demanding, which makes their deployment especially difficult for mobile terminals. Second, experimental evidence shows that CNNs often over-smooth regular textures present in images, in contrast to traditional non-local models. In this letter, we propose a solution to both issues by combining a nonlocal algorithm with a lightweight residual CNN. s solution gives full latitude to the advantages of both models. We apply this framework to two GPU implementations of classic nonlocal algorithms (NLM and BM3D) and observe a substantial gain in both cases, performing better than the state-of-the-art with low computational requirements. Our solution is between 10 and 20 times faster than CNNs with equivalent performance and attains higher PSNR. In addition the final method shows a notable gain on images containing complex textures like the ones of the MIT Moiré dataset. Yu Guo 0008, Axel Davy, Gabriele Facciolo, Jean-Michel Morel, Qiyu Jin |
IEEE Signal Process. Lett. | 1 |