VLDB 2026 Research / reviewers in the wild / expert
Huankang Guan
dblp:234/8472
· DBLP profile ↗
7ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0003-0825-8658ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Segmentation and scene understanding · 74% Video understanding and tracking · 10% Face, body and person analysis · 10% | |
| Human-computer interaction and pervasive computing
1 paper |
Wearable and physiological sensing · 100% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
glass surface detection |
1.0 | 1 | 2026 | Multi-Semantic Modeling for Glass Surface Detection in the Wild · AAAI 2026 |
Computer vision › Segmentation and scene understanding
semantic decomposition |
1.0 | 1 | 2026 | Multi-Semantic Modeling for Glass Surface Detection in the Wild · AAAI 2026 |
Computer vision › Segmentation and scene understanding › saliency detection
salient object detection |
0.9 | 1 | 2025 | A Contrastive-Learning Framework for Unsupervised Salient Object Detection · IEEE Trans. Image Process. 2025 |
Computer vision › Segmentation and scene understanding › saliency detection › salient object detection
unsupervised salient object detection |
0.9 | 1 | 2025 | A Contrastive-Learning Framework for Unsupervised Salient Object Detection · IEEE Trans. Image Process. 2025 |
Computer vision › Face, body and person analysis
human pose estimation |
0.8 | 1 | 2024 | PoseSOR: Human Pose Can Guide Our Attention · ECCV (18) 2024 |
Computer vision › Segmentation and scene understanding › saliency detection
salient object ranking |
0.8 | 1 | 2024 | SeqRank: Sequential Ranking of Salient Objects · AAAI 2024 |
Wearable and physiological sensing › eye tracking
gaze estimation |
0.8 | 1 | 2024 | PoseSOR: Human Pose Can Guide Our Attention · ECCV (18) 2024 |
Machine learning › Graph learning › graph neural network
graph convolutional network |
0.6 | 1 | 2022 | Learning Semantic Associations for Mirror Detection · CVPR 2022 |
Computer vision › Segmentation and scene understanding
mirror detection |
0.6 | 1 | 2022 | Learning Semantic Associations for Mirror Detection · CVPR 2022 |
Computer vision › Segmentation and scene understanding
scene understanding |
0.6 | 1 | 2022 | Learning Semantic Associations for Mirror Detection · CVPR 2022 |
Computer vision › Segmentation and scene understanding › instance segmentation
salient instance segmentation |
0.2 | 1 | 2024 | SeqRank: Sequential Ranking of Salient Objects · AAAI 2024 |
Methods — techniques the papers use, named apart from their topics
pose-guided attention · 1.5semantic decomposition · 1.0attention · 1.0adaptive semantic fusion · 1.0local appearance triplet loss · 0.9contrastive learning · 0.9sequential ranking · 0.8foveal-peripheral vision modeling · 0.8semantic side-path · 0.6graph convolution · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Semantic Modeling for Glass Surface Detection in the WildabstractGlass surfaces challenge object detection models as they mix the transmitted background with the reflected surrounding, creating confusing visual patterns. Previous methods relying on low-level cues (e.g., reflections and boundaries) or surrounding semantics are often unreliable in complex real-world scenarios. A glass image inherently comprises three distinct semantic components: semantics of the transmitted content, semantics of the reflected content, and semantics of the surrounding content. In this work, we observe that there is a relationship among these three types of semantics, where reflection semantics closely resembles surrounding semantics, while these two types of semantics tend to be different from the transmission semantics. For example, when on a street, we may see into a cafeteria through a glass wall, intermixed with reflection of the street, while the glass is surrounded by other street contents like shops and pedestrians, thereby creating a unique multi-semantic signature. Based on this observation, we propose the Multi-Semantic Net, MSNet, which identifies transmission, reflection, and surrounding semantics from glass images and exploits their relationships for glass surface detection. MSNet consists of two novel modules: (1) A Semantic Decomposition Module (SDM) containing Dual-Semantics Extraction Block to extract original image and reflection semantics and Semantic Elimination Block to progressively derive transmission and surrounding semantics, and (2) An Adaptive Semantic Fusion Module (ASFM) to fuse these semantic components and adaptively learn their relationships to handle varying reflection conditions. Extensive experiments demonstrate that MSNet surpasses SOTA methods on public glass detection benchmarks. Qianyu Cheng, Huankang Guan, Rynson W. H. Lau |
AAAI | 2 |
| 2025 | A Contrastive-Learning Framework for Unsupervised Salient Object DetectionabstractExisting unsupervised salient object detection (USOD) methods usually rely on low-level saliency priors, such as center and background priors, to detect salient objects, resulting in insufficient high-level semantic understanding. These low-level priors can be fragile and lead to failure when the natural images do not satisfy the prior assumptions, e.g., these methods may fail to detect those off-center salient objects causing fragmented objects in the segmentation. To address these problems, we propose to eliminate the dependency on flimsy low-level priors, and extract high-level saliency from natural images through a contrastive learning framework. To this end, we propose a Contrastive Saliency Network (CSNet), which is a prior-free and label-free saliency detector, with two novel modules: 1) a Contrastive Saliency Extraction (CSE) module to extract high-level saliency cues, by mimicking the human attention mechanism within an instance discriminative task through a contrastive learning framework, and 2) a Feature Re-Coordinate (FRC) module to recover spatial details, by calibrating high-level features with low-level features in an unsupervised fashion. In addition, we introduce a novel local appearance triplet (LAT) loss to assist the training process by encouraging similar saliency scores for regions with homogeneous appearances. Extensive experiments show that our approach is effective and outperforms state-of-the-art methods on popular SOD benchmarks. Huankang Guan, Jiaying Lin 0001, Rynson W. H. Lau |
IEEE Trans. Image Process. | 1 |
| 2024 | SeqRank: Sequential Ranking of Salient ObjectsabstractSalient Object Ranking (SOR) is the process of predicting the order of an observer's attention to objects when viewing a complex scene. Existing SOR methods primarily focus on ranking various scene objects simultaneously by exploring their spatial and semantic properties. However, their solutions of simultaneously ranking all salient objects do not align with human viewing behavior, and may result in incorrect attention shift predictions. We observe that humans view a scene through a sequential and continuous process involving a cycle of foveating to objects of interest with our foveal vision while using peripheral vision to prepare for the next fixation location. For instance, when we see a flying kite, our foveal vision captures the kite itself, while our peripheral vision can help us locate the person controlling it such that we can smoothly divert our attention to it next. By repeatedly carrying out this cycle, we can gain a thorough understanding of the entire scene. Based on this observation, we propose to model the dynamic interplay between foveal and peripheral vision to predict human attention shifts sequentially. To this end, we propose a novel SOR model, SeqRank, which reproduces foveal vision to extract high-acuity visual features for accurate salient instance segmentation while also modeling peripheral vision to select the object that is likely to grab the viewer’s attention next. By incorporating both types of vision, our model can mimic human viewing behavior better and provide a more faithful ranking among various scene objects. Most notably, our model improves the SA-SOR/MAE scores by +6.1%/-13.0% on IRSR, compared with the state-of-the-art. Extensive experiments show the superior performance of our model on the SOR benchmarks. Code is available at https://github.com/guanhuankang/SeqRank. Huankang Guan, Rynson W. H. Lau |
AAAI | 1 |
| 2024 | PoseSOR: Human Pose Can Guide Our Attention
Huankang Guan, Rynson W. H. Lau |
ECCV (18) | 1 |
| 2023 | Which Traffic Light Should You Look at? Automatically Associating Traffic Lights with Roads in High-Definition Map (Industrial Paper)abstractHigh-definition (HD) maps play an essential role in autonomous driving. However, producing HD map needs huge amount of manual annotations and is thus labor intensive and costly, which limits the widespread use of HD map. To improve the productivity and reduce cost, extensive studies have explored automation of HD map production. Existing studies primarily focus on constructing vectorized map elements from vehicle sensing images and point clouds. However, the follow-up procedure of extracting traffic semantics based on vectorized map elements is lack of study, though it is laborious and costly as well. In this paper, we focus on the automation of inferring traffic light controls for HD map production. To be specific, we aim at associating traffic lights with their controlled roads based on vectorized map data. This problem is not trivial in that: 1) the placement of traffic light contrastive to road varies considerably from scene to scene; 2) even if the placement of traffic lights and roads are similar, the road network layout has great influence on the traffic light controls. To tackle the above challenges, we propose a Heterogeneous Interaction model with Stacked Transformers (HIST) that learns representation from vectorized map elements and encodes contextual information via heterogeneous interactions among different types of map elements. We conduct extensive experiments in major cities of China to validate the efficacy of HIST. Results show HIST achieves accuracy ranging from 96.03% to 98.85% in different cities. We further deploy HIST on the HD map production line at AMAP. By incorporating a rule-based confidence system, the whole system achieves the performance with accuracy > 99.9% and automation rate > 85%, which meets the industrial level quality requirement and saves vast amount of human labor. Yitian Liao, Zan Sun, Huankang Guan, Danning Jiang, Yong Li 0008 |
SIGSPATIAL/GIS | 4 |
| 2022 | Learning Semantic Associations for Mirror DetectionabstractMirrors generally lack a consistent visual appearance, making mirror detection very challenging. Although recent works that are based on exploiting contextual contrasts and corresponding relations have achieved good results, heavily relying on contextual contrasts and corresponding relations to discover mirrors tend to fail in complex real-world scenes, where a lot of objects, e.g., doorways, may have similar features as mirrors. We observe that humans tend to place mirrors in relation to certain objects for specific functional purposes, e.g., a mirror above the sink. Inspired by this observation, we propose a model to exploit the semantic associations between the mirror and its surrounding objects for a reliable mirror localization. Our model first acquires class-specific knowledge of the surrounding objects via a semantic side-path. It then uses two novel modules to exploit semantic associations: 1) an Associations Exploration (AE) Module to extract the associations of the scene objects based on fully connected graph models, and 2) a Quadruple-Graph (QG) Module to facilitate the diffusion and aggregation of semantic association knowledge using graph convolutions. Extensive experiments show that our method outperforms the existing methods and sets the new state-of-the-art on both PMD dataset (f-measure: 0.844) and MSD dataset (f-measure: 0.889). Code is available at https://github.com/guanhuankang/Learning-Semantic-Associations-for-Mirror-Detection. Huankang Guan, Jiaying Lin 0001, Rynson W. H. Lau |
CVPR | 1 |
| 2019 | Fake News Detection via NLP is Vulnerable to Adversarial AttacksabstractNews plays a significant role in shaping people's beliefs and opinions. Fake news has always been a problem, which wasn't exposed to the mass public until the past election cycle for the 45th President of the United States. While quite a few detection methods have been proposed to combat fake news since 2015, they focus mainly on linguistic aspects of an article without any fact checking. In this paper, we argue that these models have the potential to misclassify fact-tampering fake news as well as under-written real news. Through experiments on Fakebox, a state-of-the-art fake news detector, we show that fact tampering attacks can be effective. To address these weaknesses, we argue that fact checking should be adopted in conjunction with linguistic characteristics analysis, so as to truly separate fake news from real news. A crowdsourced knowledge graph is proposed as a straw man solution to collecting timely facts about news events. Zhixuan Zhou, Huankang Guan, Meghana Moorthy Bhat, Justin Hsu |
ICAART (2) | 2 |