Gen Li 0008

dblp:28/538-8 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0001-6636-1106ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 10 since 2021
YearPublicationVenuePosition
2026 Mask2IV: Interaction-Centric Video Generation via Mask Trajectories
abstract
Generating interaction-centric videos, such as those depicting humans or robots interacting with objects, is crucial for embodied intelligence, as they provide rich and diverse visual priors for robot learning, manipulation policy training, and affordance reasoning. However, existing methods often struggle to model such complex and dynamic interactions. While recent studies show that masks can serve as effective control signals and enhance generation quality, obtaining dense and precise mask annotations remains a major challenge for real-world use. To overcome this limitation, we introduce Mask2IV, a novel framework specifically designed for interaction-centric video generation. It adopts a decoupled two-stage pipeline that first predicts plausible motion trajectories for both actor and object, then generates a video conditioned on these trajectories. This design eliminates the need for dense mask inputs from users while preserving the flexibility to manipulate the interaction process. Furthermore, Mask2IV supports versatile and intuitive control, allowing users to specify the target object of interaction and guide the motion trajectory through action descriptions or spatial position cues. To support systematic training and evaluation, we curate two benchmarks covering diverse action and object categories across both human-object interaction and robotic manipulation scenarios. Extensive experiments demonstrate that our method achieves superior visual realism and controllability compared to existing baselines.
Gen Li 0008, Jianfei Yang 0001, Laura Sevilla-Lara
AAAI1
2025 Principles of Visual Tokens for Efficient Video Understanding
abstract
Video understanding has made huge strides in recent years, relying largely on the power of transformers. As this architecture is notoriously expensive and video data is highly redundant, research into improving efficiency has become particularly relevant. Some creative solutions include token selection and merging. While most methods succeed in reducing the cost of the model and maintaining accuracy, an interesting pattern arises: most methods do not outperform the baseline of randomly discarding tokens. In this paper we take a closer look at this phenomenon and observe 5 principles of the nature of visual tokens. For example, we observe that the value of tokens follows a clear Pareto-distribution where most tokens have remarkably low value, and just a few carry most of the perceptual information. We build on these and further insights to propose a lightweight video model, LITE, that can select a small number of tokens effectively, outperforming state-of-the-art and existing baselines across datasets (Kinetics-400 and Something-Something-V2) in the challenging trade-off of computation (GFLOPs) vs accuracy. Experiments also show that LITE generalizes across datasets and even other tasks without the need for retraining.
Xinyue Hao 0001, Gen Li 0008, Shreyank N. Gowda, Robert B. Fisher, Jonathan Huang, Anurag Arnab, Laura Sevilla-Lara
ICCV2
2025 Learning Precise Affordances From Egocentric Videos for Robotic Manipulation
Gen Li 0008, Nikolaos Tsagkas, Jifei Song, Ruaridh Mon-Williams, Sethu Vijayakumar, Kun Shao, Laura Sevilla-Lara
ICCV1
2025 Dual Enhancement on 3D Vision-Language Perception for Monocular 3D Visual Grounding
abstract
Monocular 3D visual grounding is a novel task that aims to locate 3D objects in RGB images using text descriptions with explicit geometry information. Despite the inclusion of geometry details in the text, we observe that the text embeddings are sensitive to the magnitude of numerical values but largely ignore the associated measurement units. For example, simply equidistant mapping the length with unit 'meters' to 'decimeters' or 'centimeters' leads to severe performance degradation, even though the physical length remains equivalent. This observation signifies the weak 3D comprehension of pre-trained language model, which generates misguiding text features to hinder 3D perception. Therefore, we propose to enhance the 3D perception of model on text embeddings and geometry features with two simple and effective methods. Firstly, we introduce a pre-processing method named 3D-text Enhancement (3DTE), which enhances the comprehension of mapping relationships between different units by augmenting the diversity of distance descriptors in text queries. Next, we propose a Text-Guided Geometry Enhancement (TGE) module to further enhance the 3D-text information by projecting the basic text features into geometrically consistent space. These 3D-enhanced text features are then leveraged to precisely guide the attention of geometry features. We evaluate the proposed method through extensive comparisons and ablation studies on the Mono3DRefer dataset. Experimental results demonstrate substantial improvements over previous methods, achieving new state-of-the-art results with a notable accuracy gain of 11.94% in the 'Far' scenario. Our code will be made publicly available.
Min Liu 0008, Yuan Bian 0002, Zhaoyang Li 0011, Gen Li 0008, Yaonan Wang 0001
ACM Multimedia6
2024 One-Shot Open Affordance Learning with Foundation Models
abstract
We introduce One-shot Open Affordance Learning (OOAL), where a model is trained with just one example per base object category, but is expected to identify novel objects and affordances. While vision-language models excel at recognizing novel objects and scenes, they often struggle to understand finer levels of granularity such as affordances. To handle this issue, we conduct a comprehensive analysis of existing foundation models, to explore their inherent understanding of affordances and assess the potential for data-limited affordance learning. We then propose a vision-language framework with simple and effective designs that boost the alignment between visual features and affordance text embeddings. Experiments on two affordance segmentation benchmarks show that the proposed method outperforms state-of-the-art models with less than 1% of the full training data, and exhibits reasonable generalization capability on unseen objects and affordances. Project page: https://reagan1311.github.io/ooal.
Gen Li 0008, Deqing Sun, Laura Sevilla-Lara, Varun Jampani
CVPR1
2023 LOCATE: Localize and Transfer Object Parts for Weakly Supervised Affordance Grounding
abstract
Humans excel at acquiring knowledge through observation. For example, we can learn to use new tools by watching demonstrations. This skill is fundamental for intelligent systems to interact with the world. A key step to acquire this skill is to identify what part of the object affords each action, which is called affordance grounding. In this paper, we address this problem and propose a framework called LOCATE that can identify matching object parts across images, to transfer knowledge from images where an object is being used (exocentric images used for learning), to images where the object is inactive (egocentric ones used to test). To this end, we first find interaction areas and extract their feature embeddings. Then we learn to aggregate the embeddings into compact prototypes (human, object part, and background), and select the one representing the object part. Finally, we use the selected prototype to guide affordance grounding. We do this in a weakly supervised manner, learning only from image-level affordance and object labels. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods by a large margin on both seen and unseen objects.11Project page: https://reagan1311.github.io/locate.
Gen Li 0008, Varun Jampani, Deqing Sun, Laura Sevilla-Lara
CVPR1
2023 Referenceless User Controllable Semantic Image Synthesis
abstract
Despite recent progress in semantic image synthesis, complete control over image style remains a challenging problem. Existing methods require reference images to feed style information into semantic layouts, which indicates that the style is constrained by the given image. In this paper, we propose a model named RUCGAN for user controllable semantic image synthesis, which utilizes a singular color to represent the style of a specific semantic region. The proposed network achieves reference-free semantic image synthesis by injecting color as user-desired styles into each semantic layout, and is able to synthesize semantic images with unusual colors. Extensive experimental results on various challenging datasets show that the proposed method outperforms existing methods, and we further provide an interactive UI to demonstrate the advantage of our approach for style controllability. The codes and UI are available at: https://github.com/BenjaminJonghyun/RUCGAN
Jonghyun Kim 0007, Gen Li 0008, Joongkyu Kim
IJCNN2
2021 SuperStyleNet: Deep Image Synthesis with Superpixel Based Style Encoder
Jonghyun Kim 0007, Gen Li 0008, Cheolkon Jung, Joongkyu Kim
BMVC2
2021 Adaptive Prototype Learning and Allocation for Few-Shot Segmentation
abstract
Prototype learning is extensively used for few-shot segmentation. Typically, a single prototype is obtained from the support feature by averaging the global object information. However, using one prototype to represent all the information may lead to ambiguities. In this paper, we propose two novel modules, named superpixel-guided clustering (SGC) and guided prototype allocation (GPA), for multiple prototype extraction and allocation. Specifically, SGC is a parameter-free and training-free approach, which extracts more representative prototypes by aggregating similar feature vectors, while GPA is able to select matched prototypes to provide more accurate guidance. By integrating the SGC and GPA together, we propose the Adaptive Superpixel-guided Network (ASGNet), which is a lightweight model and adapts to object scale and shape variation. In addition, our network can easily generalize to k-shot segmentation with substantial improvement and no additional computational cost. In particular, our evaluations on COCO demonstrate that ASGNet surpasses the state-of-the-art method by 5% in 5-shot segmentation.1
Gen Li 0008, Varun Jampani, Laura Sevilla-Lara, Deqing Sun, Jonghyun Kim 0007, Joongkyu Kim
CVPR1
2021 Progressive Face Super-Resolution with Non-Parametric Facial Prior Enhancement
abstract
The main challenge of face super-resolution is to overcome facial distortions in an upscaling process. Recent works have utilized facial priors such as facial landmarks and component maps to generate a precise super-resolved image. However, the facial priors are estimated from the ground-truth and deep neural networks. Thus, recent works based on the facial priors are not only limited to specific datasets including the ground-truth, but also need sub-networks to extract facial priors. To solve these problems, we propose a progressive face super-resolution network with non-parametric facial prior enhancement, called as NPFNet, which extracts and highlights facial components without any tricks, such as the ground-truth and deep neural networks. The self-enhancement module facilitates our network to fully utilize facial distinct features to enhance super-resolved images with a parameter-free operation. Extensive experiments on CelebA and VGGFace2 demonstrate that the proposed method outperforms state-of-the-art face super-resolution methods in terms of visual quality and quantitative measurements.
Jonghyun Kim 0007, Gen Li 0008, Cheolkon Jung, Joongkyu Kim
ICIP2
2021 Perspective-Aware Density Regression For Crowd Counting
abstract
Scale variation, perspective distortion and severe occlusion are three main problems that affect the accuracy of crowd counting. Existing methods either adopt attention mechanism or perspective values to address these problems, but without considering them as a whole. In this paper, we advocate introducing perspective information across different density distributions to facilitate crowd estimation, and propose a novel model named Perspective-aware Density Regression Network (PDRNet) for crowd counting. Unlike previous works, PDRNet is a bilateral structure with different focuses on the features of perspective and crowd density, and it includes the information interaction module (IIM) and the similarity comparison module (SCM) to enhance the perspective-density interaction in a cross-reference manner. Specifically, IIM achieves the mutual feature guidance by swapping the global information between different branches, and SCM performs feature comparison to refine the prediction of foreground regions. With extensive experiments and competitive performances on four widely used datasets, we demonstrate the effectiveness of the proposed network.
Gen Li 0008, Joongkyu Kim, Huifang Li 0004
ICIP2
2021 Edge and identity preserving network for face super-resolution
abstract
Face super-resolution (SR) has become an indispensable function in security solutions such as video surveillance and identification system, but the distortion in facial components is a great challenge in it. Most state-of-the-art methods have utilized facial priors with deep neural networks. These methods require extra labels, longer training time, and larger computation memory. In this paper, we propose a novel Edge and Identity Preserving Network for Face SR Network, named as EIPNet, to minimize the distortion by utilizing a lightweight edge block and identity information. We present an edge block to extract perceptual edge information, and concatenate it to the original feature maps in multiple scales. This structure progressively provides edge information in reconstruction to aggregate local and global structural information. Moreover, we define an identity loss function to preserve identification of SR images. The identity loss function compares feature distributions between SR images and their ground truth to recover identities in SR images. In addition, we provide a luminance-chrominance error (LCE) to separately infer brightness and color information in SR images. The LCE method not only reduces the dependency of color information by dividing brightness and color components but also enables our network to reflect differences between SR images and their ground truth in two color spaces of RGB and YUV. The proposed method facilitates the proposed SR network to elaborately restore facial components and generate high quality 8× scaled SR images with a lightweight network structure. Furthermore, our network is able to reconstruct an 128×128 SR image with 215 fps on a GTX 1080Ti GPU. Extensive experiments demonstrate that our network qualitatively and quantitatively outperforms state-of-the-art methods on two challenging datasets: CelebA and VGGFace2.
Jonghyun Kim 0007, Gen Li 0008, Inyong Yun, Cheolkon Jung, Joongkyu Kim
Neurocomputing2
2021 Weakly-supervised temporal attention 3D network for human action recognition
Jonghyun Kim 0007, Gen Li 0008, Inyong Yun, Cheolkon Jung, Joongkyu Kim
Pattern Recognit.2
2019 DABNet: Depth-wise Asymmetric Bottleneck for Real-time Semantic Segmentation
Gen Li 0008, Joongkyu Kim
BMVC1