Wenqi Huang 0002

dblp:03/4775-2 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
17since 2021 · last 2025
0009-0004-9422-4589ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2025 FedCare: towards interactive diagnosis of federated learning systems
Tian-Ye Zhang, Haozhe Feng, Wenqi Huang 0002, Lingyu Liang, Huanming Zhang, Zexian Chen, Anthony K. H. Tung, Wei Chen 0001
Frontiers Comput. Sci.3
2024 NLSIT: A Non-Local Stereo Interaction Transformer for Stereo Image Super-Resolution
abstract
In recent years, although Transformer has been introduced into stereo image super-resolution and accomplished great advances, the long-range complementary information in stereo images hasn’t been fully utilized. In view of beneficial non-local prior knowledge in both intra-view and cross-view, we propose an efficient Non-Local Stereo Interaction Transformer (NLSIT) to exploit long-range complementary prior. NLSIT mainly consists of non-local channel interaction block (NLCIB) and non-local spatial interaction block (NLSIB). NLCIB extracts channel correlations across the two views by a channel interaction attention mechanism with linear complexity and NLSIB is devised to capture non-local spatial dependencies with locality sensitive hashing (LSH) enforcing sparsity and relevancy of the attention range. Extensive experiments demonstrate that our NLSIT outperforms most SOTA methods on several popular stereo image datasets with much fewer parameters, showing the effectiveness of the proposed framework.
Huiyun Cao, Wenqi Huang 0002, Wenming Yang
ICASSP2
2024 DEGAN: Discrimination Enhanced GAN for Perceptual-Oriented Super-Resolution
abstract
Recent years, generative adversarial networks (GANs) have gained significant prominence in single image super-resolution (SISR) tasks. This can mainly be attributed to their exceptional ability to generate intricate details. However, the instability and lack of realism in the details generated by GANs have been challenges. Existing methods mainly concentrate on improving the generator and designing complex loss functions, often overlooking the important role of discrimination. To this end, we propose our discrimination enhanced GAN (DEGAN) by improving the discriminator and simplify the discrimination task. We introduce an efficient wide activation UNet to enhance the discriminator, enabling a more comprehensive and nuanced analysis of the input image. Additionally, we introduce a texture aware mask that provides more precise guidance and alleviates the difficulty of discrimination. Our DEGAN is simple yet effective. Quantitative and visual comparisons with state-of-the-art methods on benchmark datasets demonstrate the superiority of our method.
Xiaoyu Jin, Wenqi Huang 0002, Lingyu Liang, Yang Wu 0001, Qunsheng Zeng, Ruiye Zhou, Zhuojun Cai, Jianing Shang, Wenming Yang
ICASSP2
2023 An Entity Alignment Method Based on Graph Attention Network with Pre-classification
Wenqi Huang 0002, Lingyu Liang, Yongjie Liang, Jiaxuan Hou, Xuanang Li
WISA1
2023 Retiformer: Retinex-Based Enhancement In Transformer For Low-Light Image
abstract
Transformer-based methods have shown impressive potential in many low-level vision tasks but are rarely used for low-light image enhancement (LLIE). Direct use of Transformer in LLIE will bring unnatural visual effects. This phenomenon encourages us to attempt to learn from the theory of Retinex. After trial and analysis, we finally propose Retiformer. Retiformer decomposes images into reflectance and illumination attention maps by Retinex Window Self-Attention (R-WSA). It will replace element-wise multiplication with the attention mechanism. By the R-WSA, we respectively apply a Decom-Retiformer block and an Enhance-Retiformer block at the head and tail of a Transformer-based backbone. They can decompose and align the reflection and illumination components just like RetinexNet. With this pipeline, Retiformer combines the advantages of Transformer and Retinex theory and achieves state-of-the-art performance of Retinex-based methods.
Junxiang Ruan, Xiangtao Kong, Wenqi Huang 0002, Wenming Yang
ICASSP3
2023 Spatial Correlation Fusion Network for Few-Shot Segmentation
abstract
Few-shot semantic segmentation aims to learn new knowledge rapidly with very few annotated data to segment novel classes. Recent methods follow a metric learning framework with prototypes for foreground representation [1]. However, representing support images by one or more prototypes may face problems caused by inadequate representation for segmentation, noise in complex scenes, and close semantic relation to background features. We propose a Spatial Correlation Fusion Network(SCFNet) for few-shot segmentation to address the issues. Firstly, to better capture fine-grained features, we design a Spatial Correlation Fusion module to address the loss of spatial information in support images, thus improving the performance of Few-shot segmentation. Secondly, a Prototype Contrastive Transformation(PCT) module is proposed to learn a transformation matrix for the prototype, which is capable of alleviating close semantic information and noise by adopting transformation loss. Experiments on PASCAL-5i[2] and COCO-20i[3] validate the effectiveness of our network for few-shot semantic segmentation and show our approach achieves state-of-the-art results.
Wenqi Huang 0002, Wenming Yang, Qingmin Liao
ICASSP2
2023 Global Matching-Optimization Network for Stereo Depth Estimation
abstract
Recently, iterative optimization-based approaches have gained tremendous progress in the field of stereo matching. However, it still remains a challenge to accurately estimate disparity for occlusion and textureless regions. To address this challenge, we present the Global Matching-Optimization Stereo Network (GMOStereo), which contains three components: Conv-Trans Feature Extraction Module (C-TFEM), Global Matching Module (GMM), and scene-aware disparity optimization. Before iterative optimization, attention-based GMM builds stable interdependence across distinct views. The C-TFEM, which extracts features through a two-branch network based of convolution blocks and transformer blocks, is designed to obtain global representations of features while preserving fine-grained information. The scene self-similarity adopted in disparity optimization provides supplement for matching information. Finally, a Matching-Optimization loss is designed to guide the training by imposing a direct constraint on the correlation volume. Evaluation demonstrates that GMOStereo achieves superior cross-dataset generalization performance and outperforms typical methods in the foreground and challenging regions on KITTI-2015 benchmarks.
Wenqi Huang 0002, Wenming Yang
ICASSP2
2023 Evaluation and Improvement of Interpretability for Self-Explainable Part-Prototype Networks
abstract
Part-prototype networks (e.g., ProtoPNet, ProtoTree, and ProtoPool) have attracted broad research interest for their intrinsic interpretability and comparable accuracy to non-interpretable counterparts. However, recent works find that the interpretability from prototypes is fragile, due to the semantic gap between the similarities in the feature space and that in the input space. In this work, we strive to address this challenge by making the first attempt to quantitatively and objectively evaluate the interpretability of the part-prototype networks. Specifically, we propose two evaluation metrics, termed as "consistency score" and "stability score", to evaluate the explanation consistency across images and the explanation robustness against perturbations, respectively, both of which are essential for explanations taken into practice. Furthermore, we propose an elaborated part-prototype network with a shallow-deep feature alignment (SDFA) module and a score aggregation (SA) module to improve the interpretability of prototypes. We conduct systematical evaluation experiments and provide substantial discussions to uncover the interpretability of existing part-prototype networks. Experiments on three benchmarks across nine architectures demonstrate that our model achieves significantly superior performance to the state of the art, in both the accuracy and interpretability. Our code is available at https://github.com/hqhQAQ/EvalProtoPNet.
Qihan Huang, Mengqi Xue, Wenqi Huang 0002, Haofei Zhang, Jie Song 0011, Yongcheng Jing, Mingli Song
ICCV3
2023 CKR-Calibrator: Convolution Kernel Robustness Evaluation and Calibration
Yijun Bei, Jinsong Geng, Erteng Liu, Kewei Gao, Wenqi Huang 0002, Zunlei Feng
ICONIP (5)5
2023 Attribution Guided Layerwise Knowledge Amalgamation from Graph Neural Networks
Yunzhi Hao, Yu Wang 0176, Shunyu Liu 0001, Tongya Zheng, Xingen Wang, Xinyu Wang 0001, Mingli Song, Wenqi Huang 0002, Chun Chen 0001
ICONIP (1)8
2023 Heterogeneous Graph Prototypical Networks for Few-Shot Node Classification
Yunzhi Hao, Mengfan Wang, Xingen Wang, Tongya Zheng, Xinyu Wang 0001, Wenqi Huang 0002, Chun Chen 0001
ICONIP (8)6
2023 Propheter: Prophetic Teacher Guided Long-Tailed Distribution Learning
Yongcheng Jing, Linyun Zhou, Wenqi Huang 0002, Lechao Cheng, Zunlei Feng, Mingli Song
ICONIP (4)4
2023 Disentangling Node Metric Factors for Temporal Link Prediction
Tianli Zhang, Tongya Zheng, Yuanyu Wan, Wenqi Huang 0002
ICONIP (2)5
2023 Constituent Attention for Vision Transformers
Haoling Li, Mengqi Xue, Jie Song 0011, Haofei Zhang, Wenqi Huang 0002, Lingyu Liang, Mingli Song
Comput. Vis. Image Underst.5
2023 VIS+AI: integrating visualization with artificial intelligence for efficient data analysis
abstract
Abstract Visualization and artificial intelligence (AI) are well-applied approaches to data analysis. On one hand, visualization can facilitate humans in data understanding through intuitive visual representation and interactive exploration. On the other hand, AI is able to learn from data and implement bulky tasks for humans. In complex data analysis scenarios, like epidemic traceability and city planning, humans need to understand large-scale data and make decisions, which requires complementing the strengths of both visualization and AI. Existing studies have introduced AI-assisted visualization as AI4VIS and visualization-assisted AI as VIS4AI. However, how can AI and visualization complement each other and be integrated into data analysis processes are still missing. In this paper, we define three integration levels of visualization and AI. The highest integration level is described as the framework of VIS+AI, which allows AI to learn human intelligence from interactions and communicate with humans through visual interfaces. We also summarize future directions of VIS+AI to inspire related studies.
Xumeng Wang, Ziliang Wu, Wenqi Huang 0002, Yating Wei, Zhaosong Huang, Mingliang Xu 0001, Wei Chen 0001
Frontiers Comput. Sci.3
2023 ChartNavigator: An Interactive Pattern Identification and Annotation Framework for Charts
abstract
Patterns in charts refer to interesting visual features or forms. Identifying patterns not only helps analysts understand the ‘shape’ of the data but also supports better and faster decision-making. Existing solutions for identifying patterns in charts require a large number of labeled data instances, making it intractable without user supervision. In this paper, we propose ChartNavigator, an interactive pattern identification and annotation framework for unlabeled visualization charts. ChartNavigator leverages a novel chart-sensitive deep factor model to map patterns into a low-dimensional factor representation space, and facilitates rich analysis with the derived representations. We design and implement a visual interface to support efficient identification and annotation of potential patterns in charts. Evaluations with multiple datasets show that our approach outperforms the baseline models in identifying and annotating patterns
Tian-Ye Zhang, Haozhe Feng, Wei Chen 0001, Zexian Chen, Wenting Zheng, Wenqi Huang 0002, Anthony K. H. Tung
IEEE Trans. Knowl. Data Eng.7
2022 Identification of Bird's Nest Hazard Level of Transmission Line Based on Improved Yolov5 and Location Constraints
Yang Wu 0001, Qunsheng Zeng, Wenqi Huang 0002, Lingyu Liang
PRCV (4)4
2015 Joint Object Segmentation and Depth Upsampling
abstract
With the advent of powerful ranging and visual sensors, nowadays, it is convenient to collect sparse 3-D point clouds and aligned high-resolution images. Benefitted from such convenience, this letter proposes a joint method to perform both depth assisted object-level image segmentation and image guided depth upsampling. To this end, we formulate these two tasks together as a bi-task labeling problem, defined in a Markov random field. An alternating direction method (ADM) is adopted for the joint inference, solving each sub-problem alternatively. More specifically, the sub-problem of image segmentation is solved by Graph Cuts, which attains discrete object labels efficiently. Depth upsampling is addressed via solving a linear system that recovers continuous depth values. By this joint scheme, robust object segmentation results and high-quality dense depth maps are achieved. The proposed method is applied to the challenging KITTI vision benchmark suite, as well as the Leuven dataset for validation. Comparative experiments show that our method outperforms stand-alone approaches.
Wenqi Huang 0002, Xiaojin Gong, Michael Ying Yang
IEEE Signal Process. Lett.1
2014 Road scene segmentation via fusing camera and lidar data
abstract
This paper presents an approach for pixel-wise object segmentation for road scenes based on the integration of a color image and an aligned 3D point cloud. In light of the advantage of range information in object discovery, we first produce initial object hypotheses by clustering the sparse 3D point cloud. The image pixels registered to the clustered 3D points are taken as samples to learn each object's prior knowledge. The priors are represented by Gaussian Mixture Models (GMMs) of color and 3D location information only, requiring no high-level features. We further formulate the segmentation problem within a Conditional Random Field (CRF) framework, which incorporates the learned prior models, together with hard constraints placed on the registered pixels and pairwise spatial constraints to achieve final results. Our algorithm is validated on the challenging KITTI dataset which contains diverse complicated road scenarios. Both qualitative and quantitative evaluation results show the superiority of our algorithm.
Wenqi Huang 0002, Xiaojin Gong, Zhiyu Xiang
ICRA1
2013 Integrating visual and range data for road detection
abstract
This paper presents a new method for detecting drivable road surfaces in a single image. The method takes advantage of range and visual information so that reliable results are achieved. Specifically, given LIDAR data and an aligned image, it first makes use of 3D points to estimate the ground plane and determine the horizon. Then, subsets of road and obstacle points are extracted from the 3D points based on the plane and LIDAR properties. The pixels registered to the extracted points are used to build apriori road and non-road appearance models. The road detection problem is further formulated using Markov random field whose energy function is defined based on the learned models. Constraints are also added on the energy function to place high confidence on the pixels that are registered to extracted 3D points. Extensive experiments on urban roads and highways show that our method is robust even in complicated environments.
Wenqi Huang 0002, Xiaojin Gong, Jilin Liu
ICIP1