VLDB 2026 Research / reviewers in the wild / expert
Shuo Chen 0010
dblp:00/6472-10
· DBLP profile ↗
10ranked-venue papers
6as first author
7since 2021 · last 2023
0009-0005-6092-8104ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Multi-Label Meta Weighting for Long-Tailed Dynamic Scene Graph GenerationabstractThis paper investigates the problem of scene graph generation in videos with the aim of capturing semantic relations between subjects and objects in the form of ⟨ subject, predicate, object⟩ triplets. Recognizing the predicate between subject and object pairs is imbalanced and multi-label in nature, ranging from ubiquitous interactions such as spatial relationships (e.g. in front of) to rare interactions such as twisting. In widely-used benchmarks such as Action Genome and VidOR, the imbalance ratio between the most and least frequent predicates reaches 3,218 and 3,408, respectively, surpassing even benchmarks specifically designed for long-tailed recognition. Due to the long-tailed distributions and label co-occurrences, recent state-of-the-art methods predominantly focus on the most frequently occurring predicate classes, ignoring those in the long tail. In this paper, we analyze the limitations of current approaches for scene graph generation in videos and identify a one-to-one correspondence between predicate frequency and recall performance. To make the step towards unbiased scene graph generation in videos, we introduce a multi-label meta-learning framework to deal with the biased predicate distribution. Our meta-learning framework learns a meta-weight network for each training sample over all possible label losses. We evaluate our approach on the Action Genome and VidOR benchmarks by building upon two current state-of-the-art methods for each benchmark. The experiments demonstrate that the multi-label meta-weight network improves the performance for predicates in the long tail without compromising performance for head classes, resulting in better overall performance and favorable generalizability. Code: https://github.com/shanshuo/ML-MWN. Shuo Chen 0010, Yingjun Du, Pascal Mettes, Cees Snoek |
ICMR | 1 |
| 2023 | Approximately Learning Quantum Automata
Wenjing Chu, Shuo Chen 0010, Marcello M. Bonsangue, Zenglin Shi |
TASE | 2 |
| 2022 | Non-linear Optimization Methods for Learning Regular Distributions
Wenjing Chu, Shuo Chen 0010, Marcello M. Bonsangue |
ICFEM | 2 |
| 2021 | Diagnosing Errors in Video Relation Detectors
Shuo Chen 0010, Pascal Mettes, Cees Snoek |
BMVC | 1 |
| 2021 | MVT: Multi-view Vision Transformer for 3D Object Recognition
Shuo Chen 0010, Ping Li 0001 |
BMVC | 1 |
| 2021 | Social Fabric: Tubelet Compositions for Video Relation DetectionabstractThis paper strives to classify and detect the relationship between object tubelets appearing within a video as a 〈subject-predicate-object〉 triplet. Where existing works treat object proposals or tubelets as single entities and model their relations a posteriori, we propose to classify and detect predicates for pairs of object tubelets a priori. We also propose Social Fabric: an encoding that represents a pair of object tubelets as a composition of interaction primitives. These primitives are learned over all relations, resulting in a compact representation able to localize and classify relations from the pool of co-occurring object tubelets across all timespans in a video. The encoding enables our two-stage network. In the first stage, we train Social Fabric to suggest proposals that are likely interacting. We use the Social Fabric in the second stage to simultaneously finetune and predict predicate labels for the tubelets. Experiments demonstrate the benefit of early video relation modeling, our encoding and the two-stage architecture, leading to a new state-of-the-art on two benchmarks. We also show how the encoding enables query-by-primitive-example to search for spatio-temporal video relations. Code: https://github.com/shanshuo/Social-Fabric. Shuo Chen 0010, Zenglin Shi, Pascal Mettes, Cees Snoek |
ICCV | 1 |
| 2021 | Learning Probabilistic Automata Using Residuals
Wenjing Chu, Shuo Chen 0010, Marcello M. Bonsangue |
ICTAC | 2 |
| 2020 | Interactivity Proposals for Surveillance VideosabstractThis paper introduces spatio-temporal interactivity proposals for video surveillance. Rather than focusing solely on actions performed by subjects, we explicitly include the objects that the subjects interact with. To enable interactivity proposals, we introduce the notion of interactivityness, a score that reflects the likelihood that a subject and object have an interplay. For its estimation, we propose a network containing an interactivity block and geometric encoding between subjects and objects. The network computes local interactivity likelihoods from subject and object trajectories, which we use to link intervals of high scores into spatio-temporal proposals. Experiments on an interactivity dataset with new evaluation metrics show the general benefit of interactivity proposals as well as its favorable performance compared to traditional temporal and spatio-temporal action proposals. Shuo Chen 0010, Pascal Mettes, Cees Snoek |
ICMR | 1 |
| 2019 | Interactive Exploration of Journalistic Video Footage through Multimodal Semantic MatchingabstractThis demo presents a system for journalists to explore video footage for broadcasts. Daily news broadcasts contain multiple news items that consist of many video shots and searching for relevant footage is a labor intensive task. Without the need for annotated video shots, our system extracts semantics from footage and automatically matches these semantics to query terms from the journalist. The journalist can then indicate which aspects of the query term need to be emphasized, e.g. the title or its thematic meaning. The goal of this system is to support the journalists in their search process by encouraging interaction and exploration with the system. Sarah Ibrahimi, Shuo Chen 0010, Devanshu Arya, Arthur Câmara, Yunlu Chen, Tanja Crijns, Maurits van der Goes, Thomas Mensink, Emiel van Miltenburg, Daan Odijk, William Thong, Jiaojiao Zhao, Pascal Mettes |
ACM Multimedia | 2 |
| 2016 | Visual domain adaptation using weighted subspace alignmentabstractDomain Adaptation (DA) has attracted a lot of attention in recent years. DA aims at overcoming the covariate shift in dataset and aligning multiple existing but partially related data collections. In this paper, we propose a new DA algorithm which aligns the weighted subspaces generated from source samples and target samples. The weighted subspaces of source samples are generated using weighted Principal Component Analysis (PCA). Specifically, the source samples closer to the target domain are given higher weights during the construction of subspaces, which is definitely beneficial for building an adaptable classifier. Subsequently, the weighted subspaces of source samples and the subspaces of target samples are aligned to achieve domain adaptation. Experimental results on standard datasets demonstrate the advantages of our approach over state-of-the-art DA approaches. Shuo Chen 0010, Fei Zhou 0001, Qingmin Liao |
VCIP | 1 |