EDBT 2026 Demo / reviewers in the wild / expert
Hengcan Shi
dblp:187/5695
· DBLP profile ↗
29ranked-venue papers
11as first author
20since 2021 · last 2026
0000-0002-1340-0009ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 9 first-author · 16 since 2021Artificial intelligence and machine learning · 11 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FPS: Frequency prompt synchronization for micro-expression recognition
Jiateng Liu, Hengcan Shi, Yaonan Wang 0001, Wenming Zheng |
Pattern Recognit. | 2 |
| 2026 | Glob-Diffusion: A Global Consistent Diffusion Model for Large-Scale Image GenerationabstractLarge-scale images play a crucial role in geospatial surveying, as they cover an extensively broad view and diverse objects. Due to computational limitations, existing methods rely on generating large-scale images in patches. However, the lack of global guidance in these methods often leads to significant logical errors among different patches. To address this issue, we propose a Global Consistency Diffusion model (Glob-Diffusion) for large-scale image generation. The core idea is to utilize the global consistency of small-scale images to guide the generation of large-scale images. Specifically, we introduce a Hierarchical Distributed Guidance (HDG) module that extracts patch prompts with different semantic hierarchies from small-scale images, distributedly embedding them into the generation of large-scale images to maintain global consistency across various regions. In addition, we further design a Region Guided Adapter (RGA) that dynamically optimizes the guidance strength of patch prompts by comparing differences across generated regions, effectively improving the realism of large-scale images. Our method demonstrates remarkable visual synthesis results across various natural scenes, effectively preserving global consistency in large-scale images, and also significantly enhancing the generation quality of large-scale remote sensing images. Code will be available at https://github.com/kyh433/Glob-Diffusion. Yuhan Kang, Hengcan Shi, Hao Liu 0123, Weiying Xie, Leyuan Fang, Lorenzo Bruzzone |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | DrVideo: Document Retrieval Based Long Video UnderstandingabstractMost of the existing methods for video understanding primarily focus on videos only lasting tens of seconds, with limited exploration of techniques for handling long videos. The increased number of frames in long videos poses two main challenges: difficulty in locating key information and performing long-range reasoning. Thus, we propose DrVideo, a document-retrieval-based system designed for long video understanding. Our key idea is to convert the long-video understanding problem into a long-document understanding task so as to effectively leverage the power of large language models. Specifically, DrVideo first transforms a long video into a coarse text-based long document to initially retrieve key frames and then updates the documents with the augmented key frame information. It then employs an agent-based iterative loop to continuously search for missing information and augment the document until sufficient question-related information is gathered for making the final predictions in a chain-of-thought manner. Extensive experiments on long video benchmarks confirm the effectiveness of our method. DrVideo significantly outperforms existing LLM-based state-of-the-art methods on EgoSchema benchmark (3 minutes), MovieChat-1K benchmark (10 minutes), and the long split of Video-MME benchmark (average of 44 minutes). Code is available at https://github.com/Upper9527/DrVideo. Ziyu Ma, Chenhui Gou, Hengcan Shi, Bin Sun 0001, Shutao Li 0001, Seyed Hamid Rezatofighi, Jianfei Cai 0001 |
CVPR | 3 |
| 2025 | New Multiple Sclerosis Lesion Segmentation via Calibrated Inter-patch Blending
Jin Ye 0002, Son Duy Dao, Yicheng Wu 0001, Yasmeen M. George, Thanh Nguyen-Duc, Daniel F. Schmidt, Hengcan Shi, Winston Chong, Jianfei Cai 0001 |
MICCAI (16) | 7 |
| 2025 | NaME: A Natural Micro-expression Dataset for Micro-expression Recognition in the WildabstractMicro-expressions (MEs) are involuntary facial expressions that reveal genuine emotions and have significant applications in fields such as psychology, security, and human-computer interaction. However, previous ME datasets are mainly collected in controlled laboratory environments, such as fixed views, single illumination and head movements, limited subjects and the lack of background. There are significant gaps between them and the real world. To handle this issue, we introduce a novel Natural Micro-Expression (NaME) dataset, a natural dataset collected under unconstrained real-world conditions. It encompasses (1) diverse subjects, multiple views and varying head movements ; (2) rich background information, providing a more realistic benchmark for the micro-expression recognition (MER) research. Furthermore, we propose a MER benchmark for natural environments, named MixFormer. MixFormer includes an efficient sparse attention mechanism to capture subtle facial motions from various factors, and a face-background mix of attention module to model the environment context to help MER. Extensive experiments are conducted to analyze our NaME dataset and benchmark. We believe that our dataset and benchmark will pave the way for future research in MER beyond controlled settings, facilitating the deployment of MER in practical applications. NaME is available at github.com/real-ljt/NAMEdataset. Jiateng Liu, Hengcan Shi, Haiwen Liang, Yuan Zong, Yaonan Wang 0001, Wenming Zheng |
ACM Multimedia | 2 |
| 2025 | LLMFormer: Large Language Model for Open-Vocabulary Semantic Segmentation
Hengcan Shi, Son Duy Dao, Jianfei Cai 0001 |
Int. J. Comput. Vis. | 1 |
| 2025 | RF-REN: RGB-Frequency Relation Exploration Network for Micro-Expression RecognitionabstractMicro-expression recognition (MER) has drawn increasing attention in recent years due to its ability to reveal the true feelings people want to hide. The key challenge in MER is subtle motions, which are hard to capture but crucial for MER. Existing methods usually solve this problem by magnifying all motions in the whole face and temporal sequence. However, micro-expressions (MEs) only involve a few facial areas and several temporal snippets. The all-motion magnification in previous methods cannot precisely capture these local ME motion patterns, and can easily cause spatial as well as temporal distortions, which significantly decrease the MER accuracy. In this paper, we propose an RGB-Frequency Relation Exploration Network (RF-REN), which enhances the subtle motions in refined local ME cues by exploring spatial and temporal relations in both RGB and frequency domains. Specifically, we first decompose the ME video into RGB as well as frequency domains, and conduct temporal division according to different motion stages to cover various ME local patterns. Secondly, we construct an adaptive local-global relation exploration (LGRE) module to explore the local relation cues in the spatial appearance and temporal dynamics in both domains. Finally, we propose an RGB-Frequency routing strategy to fuse the RGB and frequency cues, aiming to aggregate spatial-temporal local-global information and enhance subtle motions for MER. Extensive experiments on three databases (CASME II, SAMM and SMIC) show that the proposed model outperforms other state-of-the-art methods. Jiateng Liu, Hengcan Shi, Yaonan Wang 0001, Yuan Zong |
IEEE Signal Process. Lett. | 2 |
| 2024 | JRDB-PanoTrack: An Open-World Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human EnvironmentsabstractAutonomous robot systems have attracted increasing research attention in recent years, where environment understanding is a crucial step for robot navigation, human-robot interaction, and decision. Real-world robot systems usually collect visual data from multiple sensors and are required to recognize numerous objects and their movements in complex human-crowded settings. Traditional benchmarks, with their reliance on single sensors and limited object classes and scenarios, fail to provide the comprehensive environmental understanding robots need for accurate navigation, interaction, and decision-making. As an extension of JRDB dataset, we unveil JRDB-PanoTrack, a novel open-world panoptic segmentation and tracking benchmark, towards more comprehensive environmental perception. JRDB-PanoTrack includes (1) various data involving indoor and outdoor crowded scenes, as well as comprehensive 2D and 3D synchronized data modalities; (2) high-quality 2D spatial panoptic segmentation and temporal tracking annotations, with additional 3D label projections for further spatial understanding; (3) diverse object classes for closed- and open-world recognition benchmarks, with OSPA-based metrics for evaluation. Extensive evaluation of leading methods shows significant challenges posed by our dataset. Duy-Tho Le, Chenhui Gou, Stavya Datta, Hengcan Shi, Ian D. Reid 0001, Jianfei Cai 0001, Seyed Hamid Rezatofighi |
CVPR | 4 |
| 2024 | Diffusion Model for Robust Multi-sensor Fusion in 3D Object Detection and BEV Segmentation
Duy-Tho Le, Hengcan Shi, Jianfei Cai 0001, Seyed Hamid Rezatofighi |
ECCV (68) | 2 |
| 2024 | CA-OVS: Cluster and Adapt Mask Proposals for Open-Vocabulary Semantic Segmentation
Son Duy Dao, Hengcan Shi, Dinh Q. Phung, Jianfei Cai 0001 |
MMAsia | 2 |
| 2024 | Class Enhancement Losses With Pseudo Labels for Open-Vocabulary Semantic SegmentationabstractRecent mask proposal models have significantly improved the performance of open-vocabulary semantic segmentation. However, the use of a ‘background’ embedding during training in these methods is problematic as the resulting model tends to over-learn and assign all unseen classes as the background class instead of their correct labels. Furthermore, they ignore the semantic relationship of text embeddings, which arguably can be highly informative for open-vocabulary prediction as some classes may have close relationship with other classes. To this end, this paper proposes novel class enhancement losses to bypass the use of the ‘background’ embbedding during training, and simultaneously exploit the semantic relationship between text embeddings and mask proposals by ranking the similarity scores. To further capture the relationship between base and novel classes, we propose an effective pseudo label generation pipeline using the pretrained vision-language model. Extensive experiments on several benchmark datasets show that our method achieves overall the best performance for open-vocabulary semantic segmentation. Our method is flexible, and can also be applied to the zero-shot semantic segmentation problem. Son Duy Dao, Hengcan Shi, Dinh Q. Phung, Jianfei Cai 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Unified Open-Vocabulary Dense Visual PredictionabstractIn recent years, open-vocabulary (OV) dense visual prediction (such as OV object detection, semantic, instance and panoptic segmentations) has attracted increasing research attention. However, most of the existing approaches are task-specific, i.e., tackling each task individually. In this paper, we propose a Unified Open-Vocabulary Network (UOVN) to jointly address four common dense prediction tasks. Compared with separate models, a unified network is more desirable for diverse industrial applications. Moreover, OV dense prediction training data is relatively less. Separate networks can only leverage task-relevant training data, while a unified approach can integrate diverse data to boost individual tasks. We address two major challenges in unified OV prediction. Firstly, unlike unified methods for fixed-set predictions, OV networks are usually trained with multi-modal data. Therefore, we propose a multi-modal, multi-scale and multi-task (MMM) decoding mechanism to better exploit multi-modal information for OV recognition. Secondly, because UOVN uses data from different tasks for training, there are significant domain and task gaps. We present a UOVN training mechanism to reduce such gaps. Experiments on four datasets demonstrate the effectiveness of our UOVN. Hengcan Shi, Munawar Hayat, Jianfei Cai 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Learning Offset Probability Distribution for Accurate Object DetectionabstractObject detection combines object classification and object localization problems. Current object detection methods heavily depend on regression networks to locate objects, which are optimized with various regression loss functions to predict offsets between candidate boxes and objects. However, these regression losses are difficult to assign the appropriate penalties for samples with large offset errors, resulting in suboptimal regression networks and inaccurate object offsets. In this article, we consider object location as offset bin classification problem, and propose a distance-aware offset bin classification network optimized with multiple binary cross entropy losses to learn various offset probability distribution, including single label distribution and distance-aware label distribution. On one hand, it provides gradient contributions for different samples based on the bounded probability instead of previous incalculable offset error. On the other hand, it explores the distance correlations between discrete offset bins to facilitate network learning. Specifically, we discretize the continuous offset into a number of bins, and predict the probability of each offset bin, in which the probability should be higher for the offset bin closer to the target offsets, and vice versa. Furthermore, we propose an expectation-based offset prediction and a hierarchical focusing method to improve the precision of prediction. We conduct extensive experiments to evaluate the effectiveness of our method. In addition, our method can be conveniently and flexibly inserted into existing object detection methods, which consistently achieves a large gain based on popular anchor-based and anchor-free methods on the PASCAL VOC, MS-COCO, KITTI, and CrowdHuman datasets. Code will be released at: https://github.com/QiuHeqian/DBC . Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Hengcan Shi, Lanxiao Wang, Fanman Meng, Linfeng Xu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Transformer Scale Gate for Semantic SegmentationabstractEffectively encoding multi-scale contextual information is crucial for accurate semantic segmentation. Most of the existing transformer-based segmentation models combine features across scales without any selection, where features on sub-optimal scales may degrade segmentation outcomes. Leveraging from the inherent properties of Vision Transformers, we propose a simple yet effective module, Transformer Scale Gate (TSG), to optimally combine multi-scale features. TSG exploits cues in self and cross attentions in Vision Transformers for the scale selection. TSG is a highly flexible plug-and-play module, and can easily be incorporated with any encoder-decoder-based hierarchical vision Transformer. Extensive experiments on the Pascal Context, ADE20K and Cityscapes datasets demonstrate that the proposed feature selection strategy achieves consistent gains. Hengcan Shi, Munawar Hayat, Jianfei Cai 0001 |
CVPR | 1 |
| 2023 | CoactSeg: Learning from Heterogeneous Data for New Multiple Sclerosis Lesion Segmentation
Yicheng Wu 0001, Hengcan Shi, Bjoern Picker, Winston Chong, Jianfei Cai 0001 |
MICCAI (8) | 3 |
| 2023 | Open-Vocabulary Object Detection via Scene Graph DiscoveryabstractIn recent years, open-vocabulary (OV) object detection has attracted increasing research attention. Unlike traditional detection, which only recognizes fixed-category objects, OV detection aims to detect objects in an open category set. Previous works often leverage vision-language (VL) training data (e.g., referring grounding data) to recognize OV objects. However, they only use pairs of nouns and individual objects in VL data, while these data usually contain much more information, such as scene graphs, which are also crucial for OV detection. In this paper, we propose a novel Scene-Graph-Based Discovery Network (SGDN) that exploits scene graph cues for OV detection. Firstly, a scene-graph-based decoder (SGDecoder) including sparse scene-graph-guided attention (SSGA) is presented. It captures scene graphs and leverages them to discover OV objects. Secondly, we propose scene-graph-based prediction (SGPred), where we build a scene-graph-based offset regression (SGOR) mechanism to enable mutual enhancement between scene graph extraction and object localization. Thirdly, we design a cross-modal learning mechanism in SGPred. It takes scene graphs as bridges to improve the consistency between cross-modal embeddings for OV object classification. Experiments on COCO and LVIS demonstrate the effectiveness of our approach. Moreover, we show the ability of our model for OV scene graph detection, while previous OV scene graph generation methods cannot tackle this task. Hengcan Shi, Munawar Hayat, Jianfei Cai 0001 |
ACM Multimedia | 1 |
| 2023 | Unpaired referring expression grounding via bidirectional cross-modal matching
Hengcan Shi, Munawar Hayat, Jianfei Cai 0001 |
Neurocomputing | 1 |
| 2022 | ProposalCLIP: Unsupervised Open-Category Object Proposal Generation via Exploiting CLIP CuesabstractObject proposal generation is an important and fundamental task in computer vision. In this paper, we propose ProposalCLIP, a method towards unsupervised open-category object proposal generation. Unlike previous works which require a large number of bounding box annotations and/or can only generate proposals for limited object categories, our ProposalCLIP is able to predict proposals for a large variety of object categories without annotations, by exploiting CLIP (contrastive language-image pre-training) cues. Firstly, we analyze CLIP for unsupervised open-category proposal generation and design an objectness score based on our empirical analysis on proposal selection. Secondly, a graph-based merging module is proposed to solve the limitations of CLIP cues and merge fragmented proposals. Finally, we present a proposal regression module that extracts pseudo labels based on CLIP cues and trains a lightweight network to further refine proposals. Extensive experiments on PASCAL VOC, COCO and Visual Genome datasets show that our ProposalCLIP can better generate proposals than previous state-of-the-art methods. Our ProposalCLIP also shows benefits for downstream tasks, such as unsupervised object detection. Hengcan Shi, Munawar Hayat, Yicheng Wu 0001, Jianfei Cai 0001 |
CVPR | 1 |
| 2021 | Deep Music Retrieval for Fine-Grained Videos by Exploiting Cross-Modal-Encoded Voice-OversabstractRecently, the witness of the rapidly growing popularity of short videos on different Internet platforms has intensified the need for a background music (BGM) retrieval system. However, existing video-music retrieval methods only based on the visual modality cannot show promising performance regarding videos with fine-grained virtual contents. In this paper, we also investigate the widely added voice-overs in short videos and propose a novel framework to retrieve BGM for fine-grained short videos. In our framework, we use the self-attention (SA) and the cross-modal attention (CMA) modules to explore the intra- and the inter-relationships of different modalities respectively. For balancing the modalities, we dynamically assign different weights to the modal features via a fusion gate. For paring the query and the BGM embeddings, we introduce a triplet pseudo-label loss to constrain the semantics of the modal embeddings. As there are no existing virtual-content video-BGM retrieval datasets, we build and release two virtual-content video datasets HoK400 and CFM400. Experimental results show that our method achieves superior performance and outperforms other state-of-the-art methods with large margins. Tingtian Li, Zixun Sun, Haoruo Zhang, Ziming Wu, Hui Zhan, Yipeng Yu, Hengcan Shi |
SIGIR | 8 |
| 2021 | Query Reconstruction Network for Referring Expression Image SegmentationabstractReferring expression image segmentation aims at segmenting out the object described by a natural language query. Due to the diversity of visual content and language descriptions, it is very challenging to accurately model the correspondence between the vision and language, which inevitably produces some undesired segmentation objects from the queries. In this paper, we propose a query reconstruction network (QRN) to build more consistent corresponding relations between the language queries and object segmentation results. QRN not only generates segmentations from the queries and images but also reversely reconstructs the queries from the segmentations and the images. Through query reconstruction, QRN can confirm the vision-language consistency between the segmentations and queries. In the inference stage, for inconsistent segmentations and queries, we propose an iterative segmentation correction (ISC) method to correct them. ISC takes the difference between the reconstructed and input queries as a loss to optimize the proposed QRN. Then, the proposed QRN can generate new segmentations and queries. By iterative optimization, the segmentations can be gradually corrected. Extensive experiments on four referring expression image segmentation databases demonstrate the effectiveness of the proposed method. Hengcan Shi, Hongliang Li 0001, Qingbo Wu 0001, King Ngi Ngan |
IEEE Trans. Multim. | 1 |
| 2020 | Offset Bin Classification Network for Accurate Object DetectionabstractObject detection combines object classification and object localization problems. Most existing object detection methods usually locate objects by leveraging regression networks trained with Smooth L1loss function to predict offsets between candidate boxes and objects. However, this loss function applies the same penalties on different samples with large errors, which results in suboptimal regression networks and inaccurate offsets. In this paper, we propose an offset bin classification network optimized with cross entropy loss to predict more accurate offsets. It not only provides different penalties for different samples but also avoids the gradient explosion problem caused by the samples with large errors. Specifically, we discretize the continuous offset into a number of bins, and predict the probability of each offset bin. Furthermore, we propose an expectation-based offset prediction and a hierarchical focusing method to improve the prediction precision. Extensive experiments on the PASCAL VOC and MS-COCO datasets demonstrate the effectiveness of our proposed method. Our method outperforms the baseline methods by a large margin. Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Hengcan Shi |
CVPR | 4 |
| 2020 | Language-Aware Fine-Grained Object Representation for Referring Expression ComprehensionabstractReferring expression comprehension expects to accurately locate an object described by a language expression, which requires precise language-aware visual object representations. However, existing methods usually use rectangular object representations, such as object proposal regions and grid regions. They ignore some fine-grained object information like shapes and poses, which are often described in language expressions and important to localize objects. Additionally, rectangular object regions usually contain background contents and irrelevant foreground features, which also decrease the localization performance. To address these problems, we propose a language-aware deformable convolution model (LDC) to learn language-aware fine-grained object representations. Rather than extracting rectangular object representations, LDC adaptively samples a set of key points based on the image and language to represent objects. This type of object representations can capture more fine-grained object information (e.g., shapes and poses) and suppress noises in accordance with language and thus, boosts the object localization performance. Based on the language-aware fine-grained object representation, we next design a bidirectional interaction model (BIM) that leverages a modified co-attention mechanism to build cross-modal bidirectional interactions to further improve the language and object representations. Furthermore, we propose a hierarchical fine-grained representation network (HFRN) to learn language-aware fine-grained object representations and cross-modal bidirectional interactions at local word level and global sentence level, respectively. Our proposed method outperforms the state-of-the-art methods on the RefCOCO, RefCOCO+ and RefCOCOg datasets. Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Hengcan Shi, Taijin Zhao, King Ngi Ngan |
ACM Multimedia | 5 |
| 2020 | Hierarchical Context Features Embedding for Object DetectionabstractPixel-level segmentation has been widely used to improve object detection. Most of the existing methods refine detection features by adding the constraint of the segmentation branch or by simply embedding high-level segmentation features into detection features within the local receptive field. However, noisy segmentation features are unavoidable in real-word applications and can easily cause false positives. To address this problem, we propose a novel hierarchical context embedding module to effectively embed segmentation features into detection features. The idea of this module is to capture hierarchical context information that includes local objects or parts and nonlocal context features by learning multiple attention maps, and subsequently utilize interdependencies between features to recalibrate noisy segmentation features. Furthermore, we use this module in the proposed gated encoder-decoder network that adaptively aggregates feature maps of different resolutions based on the gate mechanism so that we can embed multiscale segmentation feature maps into detection features for more accurate detection of objects of all sizes. Experimental results demonstrate the effectiveness of the proposed method on the Pascal VOC 2012Seg dataset, the Pascal VOC dataset and the MS COCO dataset. Heqian Qiu, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Linfeng Xu 0001, King Ngi Ngan, Hengcan Shi |
IEEE Trans. Multim. | 7 |
| 2019 | Scene Parsing via Integrated Classification Model and Variance-Based RegularizationabstractScene Parsing is a challenging task in computer vision, which can be formulated as a pixel-wise classification problem. Existing deep-learning-based methods usually use one general classifier to recognize all object categories. However, the general classifier easily makes some mistakes in dealing with some confusing categories that share similar appearances or semantics. In this paper, we propose an integrated classification model and a variance-based regularization to achieve more accurate classifications. On the one hand, the integrated classification model contains multiple classifiers, not only the general classifier but also a refinement classifier to distinguish the confusing categories. On the other hand, the variance-based regularization differentiates the scores of all categories as large as possible to reduce misclassifications. Specifically, the integrated classification model includes three steps. The first is to extract the features of each pixel. Based on the features, the second step is to classify each pixel across all categories to generate a preliminary classification result. In the third step, we leverage a refinement classifier to refine the classification result, focusing on differentiating the high-preliminary-score categories. An integrated loss with the variance-based regularization is used to train the model. Extensive experiments on three common scene parsing datasets demonstrate the effectiveness of the proposed method. Hengcan Shi, Hongliang Li 0001, Qingbo Wu 0001, Zichen Song 0002 |
CVPR | 1 |
| 2018 | Key-Word-Aware Network for Referring Expression Image Segmentation
Hengcan Shi, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001 |
ECCV (6) | 1 |
| 2018 | Boosting Scene Parsing Performance via Reliable Scale PredictionabstractSegmenting objects on suitable scales is a key factor to improve the scene parsing performance. Existing methods either simply average multi-scale results or predict scales by weakly-supervised models, due to the lack of scale labels. In this paper, we propose a novel fully-supervised Scale Prediction Model. On one hand, the proposed Scale Prediction Model learns parsing scales by the strong scale supervision, which is automatically generated from the scene parsing ground truth without any extra manually annotation. On the other hand, we explore the relationship between scale and object class, and propose to use the object class information to further improve the reliability of the scale prediction. The proposed Scale Prediction Model improves 23.1%, 20.1% and 29.3% scale prediction accuracies on the NYU Depth v2, PASCAL-Context and SIFT Flow datasets, respectively. Based on the Scale Prediction Model, we design a Scale Parsing Net (SPNet) for scene parsing, which segments each object on the scale predicted by the Scale Prediction Model. Moreover, SPNet leverages the intermediate result (i.e., the object class) to refine the parsing results. The experiment results show that SPNet outperforms many state-of-the-art methods on multiple scene parsing datasets. Hengcan Shi, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, King Ngi Ngan |
ACM Multimedia | 1 |
| 2018 | Hierarchical Parsing Net: Semantic Scene Parsing From Global Scene to ObjectsabstractThis paper proposes a novel Hierarchical Parsing Net (HPN) for semantic scene parsing. Unlike previous methods, which separately classify each object, HPN leverages global scene semantic information and the context among multiple objects to enhance scene parsing. On the one hand, HPN uses the global scene category to constrain the semantic consistency between the scene and each object. On the other hand, the context among all objects is also modeled to avoid incompatible object predictions. Specifically, HPN consists of four steps. In the first step, we extract scene and local appearance features. Based on these appearance features, the second step is to encode a contextual feature for each object, which models both the scene-object context (the context between the scene and each object) and the interobject context (the context among different objects). In the third step, we classify the global scene and then use the scene classification loss and a backpropagation algorithm to constrain the scene feature encoding. In the fourth step, a label map for scene parsing is generated from the local appearance and contextual features. Our model outperforms many state-of-the-art deep scene parsing networks on five scene parsing databases. Hengcan Shi, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Linfeng Xu 0001, King Ngi Ngan |
IEEE Trans. Multim. | 1 |
| 2017 | Improving object proposals with top-down cues
Wei Li 0110, Hongliang Li 0001, Bing Luo 0003, Hengcan Shi, Qingbo Wu 0001, King Ngi Ngan |
Signal Process. Image Commun. | 4 |
| 2016 | Person re-identification based on multi-region-set ensembles
Wei Li 0110, Chao Huang 0003, Bing Luo 0003, Fanman Meng, Tiecheng Song, Hengcan Shi |
J. Vis. Commun. Image Represent. | 6 |