EDBT 2026 Demo / reviewers in the wild / expert
Suchen Wang
dblp:213/4076
· DBLP profile ↗
13ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-4086-0713ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Action Distribution Flow for Open-Set Temporal Action SegmentationabstractIn this paper, we tackle the open-set temporal action segmentation task, which aims to identify unknown frames while ensuring accurate segmentation of known actions in the temporal domain. Existing open-set methods struggle with identifying unknown frames due to their indistinguishability against ambiguous known frames during action transitions, resulting in significant performance degradation. To address this, we propose the action distribution flow, which models transitions between action sequences to capture the inherent feature discrepancies between unknown and known frames. Specifically, our method first models the distributions of known actions using the training data, and then interpolates these distributions along the optimal transport path for consecutive actions in the testing videos. By evaluating the likelihood of testing frames against the modeled action distribution flow, our approach effectively identifies unknown frames without requiring additional training or prior knowledge of the unknown data. Extensive experiments on open-set versions of the GTEA, 50Salads, and Breakfast datasets demonstrate the superiority of the proposed method across all evaluation metrics. Runzhong Zhang, Fengrui Tian, Yueqi Duan, Ziwei Wang 0010, Weipeng Hu, Peijun Bao, Suchen Wang, Yap-Peng Tan |
IEEE Trans. Image Process. | 9 |
| 2025 | Boundary Voting Network for Ambiguity-Aware Timestamp-Supervised Action SegmentationabstractTimestamp-supervised action segmentation aims to segment and classify actions in untrimmed videos with a random frame annotated per action. Precisely localizing action boundaries from timestamp annotations is crucial for this setting, as it enables generating framewise pseudo-labels and applying the well-explored fully-supervised training. However, prevailing methods struggle with intrinsic uncertainty in boundary localization due to less discriminative features in action-transiting regions. This imprecise boundary estimation significantly reduces the stability and reliability of the generated pseudo-labels in ambiguous action-transiting regions, consequently resulting in performance deterioration of the trained segmentation models. In our paper, we introduce the boundary voting network that mitigates feature ambiguity by hierarchically propagating video-level global prior knowledge into local action-transiting regions. By generating key action representations as votes throughout the video and targeting action-transiting regions, all votes collaboratively contribute to action-transiting feature enhancement and boundary localization refinement. Extensive experiments demonstrate the effectiveness of our method on GTEA, 50Salads, and Breakfast datasets. Runzhong Zhang, Yueqi Duan, Weipeng Hu, Suchen Wang, Yap-Peng Tan |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Top-down framework for weakly-supervised grounded image captioning
Suchen Wang, Kim-Hui Yap, Yi Wang 0068 |
Knowl. Based Syst. | 2 |
| 2023 | HOI-aware Adaptive Network for Weakly-supervised Action SegmentationabstractIn this paper, we propose an HOI-aware adaptive network named AdaAct for weakly-supervised action segmentation. Most existing methods learn a fixed network to predict the action of each frame with the neighboring frames. However, this would result in ambiguity when estimating similar actions, such as pouring juice and pouring coffee. To address this, we aim to exploit temporally global but spatially local human-object interactions (HOI) as video-level prior knowledge for action segmentation. The long-term HOI sequence provides crucial contextual information to distinguish ambiguous actions, where our network dynamically adapts to the given HOI sequence at test time. More specifically, we first design a video HOI encoder that extracts, selects, and integrates the most representative HOI throughout the video. Then, we propose a two-branch HyperNetwork to learn an adaptive temporal encoder, which automatically adjusts the parameters based on the HOI information of various videos on the fly. Extensive experiments on two widely-used datasets including Breakfast and 50Salads demonstrate the effectiveness of our method under different evaluation metrics. Runzhong Zhang, Suchen Wang, Yueqi Duan, Yansong Tang, Yue Zhang 0065, Yap-Peng Tan |
IJCAI | 2 |
| 2023 | POAR: Towards Open Vocabulary Pedestrian Attribute RecognitionabstractPedestrian attribute recognition (PAR) aims to predict the attributes of a target pedestrian. Recent methods often address the PAR problem by training a multi-label classifier with predefined attribute classes, but they can hardly exhaust all possible pedestrian attributes in the real world. To tackle this problem, we propose a novel Pedestrian Open-Attribute Recognition (POAR) approach by formulating the problem as a task of image-text search. Our approach employs a Transformer-based Encoder with a Masking Strategy (TEMS) to focus on the attributes of specific pedestrian parts (e.g., head, upper body, lower body, feet, etc.), and introduces a set of attribute tokens to encode the corresponding attributes into visual embeddings. Each attribute category is described as a natural language sentence and encoded by the text encoder. Then, we compute the similarity between the visual and text embeddings to find the best attribute descriptions for the input images. To handle multiple attributes of a single pedestrian, we propose a Many-To-Many Contrastive (MTMC) loss with masked tokens. In addition, we propose a Grouped Knowledge Distillation (GKD) method to minimize the disparity between visual embeddings and unseen attribute text embeddings. We evaluate our proposed method on three PAR datasets with an open-attribute setting. The results demonstrate the effectiveness of our method as a strong baseline for the POAR task. Our code is available at https://github.com/IvyYZ/POAR. Yue Zhang 0065, Suchen Wang, Shichao Kan, Zhenyu Weng, Yi-Gang Cen, Yap-Peng Tan |
ACM Multimedia | 2 |
| 2023 | VLT: Vision-Language Transformer and Query Generation for Referring SegmentationabstractWe propose a Vision-Language Transformer (VLT) framework for referring segmentation to facilitate deep interactions among multi-modal information and enhance the holistic understanding to vision-language features. There are different ways to understand the dynamic emphasis of a language expression, especially when interacting with the image. However, the learned queries in existing transformer works are fixed after training, which cannot cope with the randomness and huge diversity of the language expressions. To address this issue, we propose a Query Generation Module, which dynamically produces multiple sets of input-specific queries to represent the diverse comprehensions of language expression. To find the best among these diverse comprehensions, so as to generate a better mask, we propose a Query Balance Module to selectively fuse the corresponding responses of the set of queries. Furthermore, to enhance the model's ability in dealing with diverse language expressions, we consider inter-sample learning to explicitly endow the model with knowledge of understanding different language expressions to the same object. We introduce masked contrastive learning to narrow down the features of different expressions for the same target object while distinguishing the features of different objects. The proposed approach is lightweight and achieves new state-of-the-art referring segmentation results consistently on five datasets. Henghui Ding, Chang Liu 0072, Suchen Wang, Xudong Jiang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Learning Transferable Human-Object Interaction Detector with Natural Language SupervisionabstractIt is difficult to construct a data collection including all possible combinations of human actions and interacting objects due to the combinatorial nature of human-object interactions (HOI). In this work, we aim to develop a transferable HOI detector for unseen interactions. Existing HOI detectors often treat interactions as discrete labels and learn a classifier according to a predetermined category space. This is inherently inapt for detecting unseen interactions which are out of the predefined categories. Conversely, we treat independent HOI labels as the natural language supervision of interactions and embed them into a joint visual-and-text space to capture their correlations. More specifically, we propose a new HOI visual encoder to detect the interacting humans and objects, and map them to a joint feature space to perform interaction recognition. Our visual encoder is instantiated as a Vision Transformer with new learnable HOI tokens and a sequence parser to generate unique HOI predictions. It distills and leverages the transferable knowledge from the pretrained CLIP model to perform the zero-shot interaction detection. Experiments on two datasets, SWIG-HOI and HICO-DET, validate that our proposed method can achieve a notable mAP improvement on detecting both seen and unseen HOIs. Our code is available at https://github.com/scwangdyd/promting_hoi. Suchen Wang, Yueqi Duan, Henghui Ding, Yap-Peng Tan, Kim-Hui Yap, Junsong Yuan 0001 |
CVPR | 1 |
| 2022 | Attribute Conditioned Fashion Image CaptioningabstractFashion image captioning aims to automatically generate product descriptions for fashion items. Existing fashion image captioning models predict a fixed caption for a particular fashion item once deployed. Differently, we explore a controllable way of fashion image captioning by taking the semantic attributes as the control signal. We propose a new multi-modal fashion image captioning method that allows the users to specify a few semantic attributes to guide the caption generation to suit unique preferences. To study the problem, we clean, filter, and assemble a new fashion image caption dataset called FACAD170K from the current FACAD dataset to facilitate learning. We investigate the effectiveness of the proposed approach on our assembled FACAD170K dataset. The results demonstrate that our method can outperform existing fashion image captioning models as well as conventional captioning methods. Kim-Hui Yap, Suchen Wang |
ICIP | 3 |
| 2021 | Vision-Language Transformer and Query Generation for Referring SegmentationabstractIn this work, we address the challenging task of referring segmentation. The query expression in referring segmentation typically indicates the target object by describing its relationship with others. Therefore, to find the target one among all instances in the image, the model must have a holistic understanding of the whole image. To achieve this, we reformulate referring segmentation as a direct attention problem: finding the region in the image where the query language expression is most attended to. We introduce transformer and multi-head attention to build a network with an encoder-decoder attention mechanism architecture that "queries" the given image with the language expression. Furthermore, we propose a Query Generation Module, which produces multiple sets of queries with different attention weights that represent the diversified comprehensions of the language expression from different aspects. At the same time, to find the best way from these diversified comprehensions based on visual clues, we further propose a Query Balance Module to adaptively select the output features of these queries for a better mask generation. Without bells and whistles, our approach is light-weight and achieves new state-of-the-art performance consistently on three referring segmentation datasets, RefCOCO, RefCOCO+, and G-Ref. Our code is available at https://github.com/henghuiding/Vision-Language-Transformer. Henghui Ding, Chang Liu 0072, Suchen Wang, Xudong Jiang 0001 |
ICCV | 3 |
| 2021 | Discovering Human Interactions with Large-Vocabulary Objects via Query and Multi-Scale DetectionabstractIn this work, we study the problem of human-object interaction (HOI) detection with large vocabulary object categories. Previous HOI studies are mainly conducted in the regime of limit object categories (e.g., 80 categories). Their solutions may face new difficulties in both object detection and interaction classification due to the increasing diversity of objects (e.g., 1000 categories). Different from previous methods, we formulate the HOI detection as a query problem. We propose a unified model to jointly discover the target objects and predict the corresponding interactions based on the human queries, thereby eliminating the need of using generic object detectors, extra steps to associate human-object instances, and multi-stream interaction recognition. This is achieved by a repurposed Transformer unit and a novel cascade detection over multi-scale feature maps. We observe that such a highly-coupled solution brings benefits for both object detection and interaction classification in a large vocabulary setting. To study the new challenges of the large vocabulary HOI detection, we assemble two datasets from the publicly available SWiG and 100 Days of Hands datasets. Experiments on these datasets validate that our proposed method can achieve a notable mAP improvement on HOI detection with a faster inference speed than existing one-stage HOI detectors. Our code is available at https://github.com/scwangdyd/large_vocabulary_hoi_detection. Suchen Wang, Kim-Hui Yap, Henghui Ding, Jiyan Wu, Junsong Yuan 0001, Yap-Peng Tan |
ICCV | 1 |
| 2020 | Discovering Human Interactions With Novel Objects via Zero-Shot LearningabstractWe aim to detect human interactions with novel objects through zero-shot learning. Different from previous works, we allow unseen object categories by using its semantic word embedding. To do so, we design a human-object region proposal network specifically for the human-object interaction detection task. The core idea is to leverage human visual clues to localize objects which are interacting with humans. We show that our proposed model can outperform existing methods on detecting interacting objects, and generalize well to novel objects. To recognize objects from unseen categories, we devise a zero-shot classification module upon the classifier of seen categories. It utilizes the classifier logits for seen categories to estimate a vector in the semantic space, and then performs nearest search to find the closest unseen category. We validate our method on V-COCO and HICO-DET datasets, and obtain superior results on detecting human interactions with both seen and unseen objects. Suchen Wang, Kim-Hui Yap, Junsong Yuan 0001, Yap-Peng Tan |
CVPR | 1 |
| 2019 | Joint Representative Selection and Feature Learning: A Semi-Supervised ApproachabstractIn this paper, we propose a semi-supervised approach for representative selection, which finds a small set of representatives that can well summarize a large data collection. Given labeled source data and big unlabeled target data, we aim to find representatives in the target data, which can not only represent and associate data points belonging to each labeled category, but also discover novel categories in the target data, if any. To leverage labeled source data, we guide representative selection from labeled source to unlabeled target. We propose a joint optimization framework which alternately optimizes (1) representative selection in the target data and (2) discriminative feature learning from both the source and the target for better representative selection. Experiments on image and video datasets demonstrate that our proposed approach not only finds better representatives, but also can discover novel categories in the target data that are not in the source. Suchen Wang, Jingjing Meng, Junsong Yuan 0001, Yap-Peng Tan |
CVPR | 1 |
| 2018 | Video Summarization Via Multiview Representative SelectionabstractVideo contents are inherently heterogeneous. To exploit different feature modalities in a diverse video collection for video summarization, we propose to formulate the task as a multiview representative selection problem. The goal is to select visual elements that are representative of a video consistently across different views (i.e., feature modalities). We present in this paper the multiview sparse dictionary selection with centroid co-regularization method, which optimizes the representative selection in each view, and enforces that the view-specific selections to be similar by regularizing them towards a consensus selection. We also introduce a diversity regularizer to favor a selection of diverse representatives. The problem can be efficiently solved by an alternating minimizing optimization with the fast iterative shrinkage thresholding algorithm. Experiments on synthetic data and benchmark video datasets validate the effectiveness of the proposed approach for video summarization, in comparison with other video summarization methods and representative selection methods such as K-medoids, sparse dictionary selection, and multiview clustering. Jingjing Meng, Suchen Wang, Hongxing Wang 0001, Junsong Yuan 0001, Yap-Peng Tan |
IEEE Trans. Image Process. | 2 |