Charig Yang

dblp:279/3612 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
5since 2021 · last 2025
0009-0003-7044-1901ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Video understanding and tracking · 58% 3D vision · 27% Image recognition and object detection · 11%
Human-computer interaction and pervasive computing
1 paper
Wearable and physiological sensing · 100%

Topics — the 8 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
activity recognition
0.912025
Reading Recognition in the Wild · NeurIPS 2025
Computer vision › 3D vision
egocentric vision
0.912025
Reading Recognition in the Wild · NeurIPS 2025
Wearable and physiological sensing › wearable camera › egocentric vision
egocentric sensing
0.912025
Reading Recognition in the Wild · NeurIPS 2025
Computer vision › 3D vision
pose estimation
0.612022
It's About Time: Analog Clock Reading in the Wild · CVPR 2022
Computer vision › Video understanding and tracking
motion segmentation
0.512021
Self-supervised Video Object Segmentation by Motion Grouping · ICCV 2021
Computer vision › Video understanding and tracking › video object segmentation
self-supervised video object segmentation
0.512021
Self-supervised Video Object Segmentation by Motion Grouping · ICCV 2021
Computer vision › Video understanding and tracking
video object segmentation
0.512021
Self-supervised Video Object Segmentation by Motion Grouping · ICCV 2021
Wearable and physiological sensing › wearable display
smart glasses
0.312025
Reading Recognition in the Wild · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

transformer · 2.2head pose · 1.7eye gaze · 1.7spatial transformer network · 1.1pseudo-labeling · 1.1self-supervised learning · 0.8optical flow · 0.5
YearPublicationVenuePosition
2025 Reading Recognition in the Wild
abstract
To enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world, including during reading. In this paper, we introduce a new task of reading recognition to determine when the user is reading. We first introduce the first-of-its-kind large-scale multimodal Reading in the Wild dataset, containing 100 hours of reading and non-reading videos in diverse and realistic scenarios. We then identify three modalities (egocentric RGB, eye gaze, head pose) that can be used to solve the task, and present a flexible transformer model that performs the task using these modalities, either individually or combined. We show that these modalities are relevant and complementary to the task, and investigate how to efficiently and effectively encode each modality. Additionally, we show the usefulness of this dataset towards classifying types of reading, extending current reading understanding studies conducted in constrained settings to larger scale, diversity and realism. Code, model, and data will be public.
Charig Yang, Samiul Alam, Shakhrul Iman Siam, Michael J. Proulx, Lambert Mathias, Kiran K. Somasundaram, Luis Pesqueira, James Fort, Sheroze Sheriffdeen, Omkar M. Parkhi, Carl Yuheng Ren, Mi Zhang 0002, Yuning Chai, Richard A. Newcombe, Hyo Jin Kim 0004
NeurIPS1
2024 Moving Object Segmentation: All You Need is SAM (and Flow)
Junyu Xie, Charig Yang, Weidi Xie, Andrew Zisserman
ACCV (10)2
2024 Made to Order: Discovering Monotonic Temporal Changes via Self-supervised Video Ordering
Charig Yang, Weidi Xie, Andrew Zisserman
ECCV (74)1
2022 It's About Time: Analog Clock Reading in the Wild
abstract
In this paper, we present a framework for reading analog clocks in natural images or videos. Specifically, we make the following contributions: First, we create a scalable pipeline for generating synthetic clocks, significantly reducing the requirements for the labour-intensive annotations; Second, we introduce a clock recognition architecture based on spatial transformer networks (STN), which is trained end-to-end for clock alignment and recognition. We show that the model trained on the proposed synthetic dataset generalises towards real clocks with good accuracy, advocating a Sim2Real training regime; Third, to further reduce the gap between simulation and real data, we leverage the special property of “time”, i.e. uniformity, to generate reliable pseudo-labels on real unlabelled clock videos, and show that training on these videos offers further improvements while still requiring zero manual annotations. Lastly, we introduce three benchmark datasets based on COCO, Open Images, and The Clock movie, with full annotations for time, accurate to the minute.
Charig Yang, Weidi Xie, Andrew Zisserman
CVPR1
2021 Self-supervised Video Object Segmentation by Motion Grouping
abstract
Animals have evolved highly functional visual systems to understand motion, assisting perception even under complex environments. In this paper, we work towards developing a computer vision system able to segment objects by exploiting motion cues, i.e. motion segmentation. To achieve this, we introduce a simple variant of the Transformer to segment optical flow frames into primary objects and the background, which can be trained in a self-supervised manner, i.e. without using any manual annotations. Despite using only optical flow, and no appearance information, as input, our approach achieves superior results compared to previous state-of-the-art self-supervised methods on public benchmarks (DAVIS2016, SegTrackv2, FBMS59), while being an order of magnitude faster. On a challenging camouflage dataset (MoCA), we significantly outperform other self-supervised approaches, and are competitive with the top supervised approach, highlighting the importance of motion cues and the potential bias towards appearance in existing video segmentation models.
Charig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman, Weidi Xie
ICCV1
2020 Betrayed by Motion: Camouflaged Object Discovery via Motion Segmentation
Hala Lamdouar, Charig Yang, Weidi Xie, Andrew Zisserman
ACCV (2)2