VLDB 2026 Research / reviewers in the wild / expert
Lan Yang 0014
dblp:52/1313-14
· DBLP profile ↗
13ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-7672-3841ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic MusicabstractExisting state-of-the-art symbolic music generation models represent symbolic music as a sequence of attribute tokens with fixed unidirectional dependencies.However, from the perspective of music theory, the attributes of a musical note are inherently a set rather than a sequence.Building on this insight, we propose Amadeus, a novel symbolic music generation framework that adopts a two-level architecture: an autoregressive model for note sequences and a bidirectional discrete diffusion model for note attributes.This design enables flexible attribute control and adjustable decoding speed during inference.To further enhance sequential modeling, we introduce the Conditional Information Enhancement Module (CIEM).We also constructed AMD (Amadeus MIDI Dataset)-the largest open-source symbolic music dataset to date-supporting both pre-training and finetuning.We trained two models of different scales, Amadeus and Amadeus-M, and conducted extensive experiments, demonstrating substantial improvements over state-of-the-art methods across both objective and subjective metrics. Hongju Su, Ke Li 0004, Lan Yang 0014, Honggang Zhang 0002, Yi-Zhe Song |
ACL (1) | 3 |
| 2026 | SACG++: Complex Sketch Generation via Representation-Enhanced Scale-Adaptive Classifier Guidance
Ke Li 0004, Jijin Hu, Lan Yang 0014, Yonggang Qi, Yi-Zhe Song |
Int. J. Comput. Vis. | 4 |
| 2026 | ST-VA-AR: Learning velocity-aware action representations with mixture of spatiotemporal attention
Jiangning Wei, Ke Li 0004, Lan Yang 0014, Dandan Xiao, Jun Liu 0014 |
Pattern Recognit. | 4 |
| 2025 | VersaGen: Unleashing Versatile Visual Control for Text-to-Image SynthesisabstractDespite the rapid advancements in text-to-image (T2I) synthesis, enabling precise visual control remains a significant challenge. Existing works attempted to incorporate multi-facet controls (text and sketch), aiming to enhance the creative control over generated images. However, our pilot study reveals that the expressive power of humans far surpasses the capabilities of current methods. Users desire a more versatile approach that can accommodate their diverse creative intents, ranging from controlling individual subjects to manipulating the entire scene composition. We present VersaGen, a generative AI agent that enables versatile visual control in T2I synthesis. VersaGen admits four types of visual controls: i) single visual subject; ii) multiple visual subjects; iii) scene background; iv) any combination of the three above or merely no control at all. We train an adaptor upon a frozen T2I model to accommodate the visual information into the text-dominated diffusion process. We introduce three optimization strategies during the inference phase of VersaGen to improve generation results and enhance user experience. Comprehensive experiments on COCO and Sketchy validate the effectiveness and flexibility of VersaGen, as evidenced by both qualitative and quantitative results. Lan Yang 0014, Yonggang Qi, Honggang Zhang 0002, Kaiyue Pang, Ke Li 0004, Yi-Zhe Song |
AAAI | 2 |
| 2025 | V-Oracle: Making Progressive Reasoning in Deciphering Oracle Bones for You and MeabstractOracle Bone Script (OBS) is a vital treasure of human civilization, rich in insights from ancient societies. However, the evolution of written language over millennia complicates its decipherment. In this paper, we propose V-Oracle, an innovative framework that utilizes Large Multi-modal Models (LMMs) for interpreting OBS. V-Oracle applies principles of pictographic character formation and frames the task as a visual question-answering (VQA) problem, establishing a multi-step reasoning chain. It proposes a multi-dimensional data augmentation for synthesizing high-quality OBS samples, and also implements a multi-phase oracle alignment tuning to improve LMMs’ visual reasoning capabilities. Moreover, to bridge the evaluation gap in the OBS field, we further introduce Oracle-Bench, a comprehensive benchmark that emphasizes process-oriented assessment and incorporates both standard and out-of-distribution setups for realistic evaluation. Extensive experimental results can demonstrate the effectiveness of our method in providing quantitative analyses and superior deciphering capability. Runqi Qiao, Qiuna Tan, Guanting Dong 0001, MinhuiWu MinhuiWu, Jiapeng Wang 0005, Zhuoma Gongque, Yadong Xue, Zhimin Bao, Lan Yang 0014, Chen Li 0031, Honggang Zhang 0002 |
ACL (1) | 13 |
| 2025 | Parameter-Efficient Adaptation of Vision-Language Models for Free-Hand Sketch RecognitionabstractHow to prompt a foundation model like CLIP towards a sketch expert is the question we seek to answer in this paper. Debates on the best way to prompt have been intense and divided, however converged on one particular point that of modelling prompt learning as context token optimisation. This paper scrutinises such technical route for sketch and argues the challenge is more than a stereotyped ask from context change. In particular, we pin down the problem to the dramatic cross-modality gap between sketch and the photo-centric visual world formed within CLIP. We first show through a pilot study that relocating a sketched object to a different spatial locality can significantly improve zero-shot CLIP performance on sketch. Our core contribution is then to regard spatial misalignment as the key to explaining poor sketch adaptation in CLIP prompts – that a sketched object does not reside in a place as if it were part of the scene compositions of photo. Methodologically, we leverage a lightweight network that explicitly allows differentiable spatial manipulation of sketch data and design regulatory self-supervised signals to encourage proper convergence. We showcase consistent complementary power of this simple approach by building on top of 10 existing contemporary prompting methods on the sketch recognition task. For example, we outperform the strong prompting baseline CoOp by 2.57%, MaPle by 4.83% and AdaptFormer by 5.07%. Notably, the latter two beat the traditional full parameter fine-tuning (82.98%83.39% vs. 81.51%), and does so with less than 1% of the total training parameters. Lan Yang 0014, Kaiyue Pang, Honggang Zhang 0002, Yi-Zhe Song |
VCIP | 2 |
| 2024 | Making Visual Sense of Oracle Bones for You and MeabstractVisual perception evolves over time. This is particularly the case of oracle bone scripts, where visual glyphs seem intuitive to people from distant past prove difficult to be understood in contemporary eyes. While semantic correspon-dence of an oracle can be found via a dictionary lookup, this proves to be not enough for public viewers to connect the dots, i.e., why does this oracle mean that? Common solution relies on a laborious curation process to collect visual guide for each oracle (Fig. 1), which hinges on the case-by-case effort and taste of curators. This paper delves into one natural follow-up question: can AI take over? Begin with a comprehensive human study, we show par-ticipants could indeed make better sense of an oracle glyph subjected to a proper visual guide and its efficacy can be approximated via a novel metric termed TransOV (Trans-ferable Oracle Visuals). We then define a new conditional visual generation task based on an oracle glyph and its se-mantic meaning and importantly approach it by circumventing any form of model training in the presence of fatal lack of oracle data. At its heart is to leverage foundation model like GPT-4V to reason about the visual cues hidden inside an oracle and take advantage of an existing text-to-image model for final visual guide generation. Extensive empirical evidence shows our AI-enabled visual guides achieve signif-icantly comparable TransOV performance compared with those collected under manual efforts. Finally, we demon-strate the versatility of our system under a more complex setting, where it is required to work alongside with an AI image denoiser to cope with raw oracle scan image inputs (cf processed clean oracle glyphs). Code is available at https://github.com/RQ-Lab/OBS-Visual. Runqi Qiao, Lan Yang 0014, Kaiyue Pang, Honggang Zhang 0002 |
CVPR | 2 |
| 2024 | Wired Perspectives: Multi-View Wire Art Embraces Generative AIabstractCreating multi-view wire art (MVWA), a static 3D sculpture with diverse interpretations from different viewpoints, is a complex task even for skilled artists. In response, we present DreamWire, an AI system enabling everyone to craft MVWA easily. Users express their vision through text prompts or scribbles, freeing them from intricate 3D wire organisation. Our approach synergises 3D Bézier curves, Prim's algorithm, and knowledge distillation from diffusion models or their variants (e.g., ControlNet). This blend enables the system to represent 3D wire art, ensuring spatial continuity and overcoming data scarcity. Extensive evaluation and analysis are conducted to shed insight on the inner workings of the proposed system, including the trade-off between connectivity and visual aesthetics. Zhiyu Qu, Lan Yang 0014, Honggang Zhang 0002, Tao Xiang 0002, Kaiyue Pang, Yi-Zhe Song |
CVPR | 2 |
| 2024 | Annotation-Free Human Sketch Quality Assessment
Lan Yang 0014, Kaiyue Pang, Honggang Zhang 0002, Yi-Zhe Song |
Int. J. Comput. Vis. | 1 |
| 2022 | Finding Badly Drawn BunniesabstractAs lovely as bunnies are, your sketched version would probably not do it justice (Fig. 1). This paper recognises this very problem and studies sketch quality measurement for the first time - letting you find these badly drawn ones. Our key discovery lies in exploiting the magnitude ($L$2norm) of a sketch feature as a quantitative quality metric. We propose Geometry-Aware Classification Layer (GACL), a generic method that makes feature-magnitude-as-quality-metric possible and importantly does it without the need for specific quality annotations from humans. GACL sees feature magnitude and recognisability learning as a dual task, which can be simultaneously optimised under a neat crossentropy classification loss. GACL is lightweight with theoretic guarantees and enjoys a nice geometric interpretation to reason its success. We confirm consistent quality agreements between our GACL-induced metric and human perception through a carefully designed human study. Notably, we demonstrate three practical sketch applications enabled for the first time using our quantitative quality metric. Lan Yang 0014, Kaiyue Pang, Honggang Zhang 0002, Yi-Zhe Song |
CVPR | 1 |
| 2021 | SketchAA: Abstract Representation for Abstract SketchesabstractWhat makes free-hand sketches appealing for humans lies with its capability as a universal tool to depict the visual world. Such flexibility at human ease, however, introduces abstract renderings that pose unique challenges to computer vision models. In this paper, we propose a purpose-made sketch representation for human sketches. The key intuition is that such representation should be abstract at design, so to accommodate the abstract nature of sketches. This is achieved by interpreting sketch abstraction on two levels: appearance and structure. We abstract sketch structure as a pre-defined coarse-to-fine visual block hierarchy, and average visual features within each block to model appearance abstraction. We then discuss three general strategies on how to exploit feature synergy across different levels of this abstraction hierarchy. The superiority of explicitly abstracting sketch representation is empirically validated on a number of sketch analysis tasks, including sketch recognition, fine-grained sketch-based image retrieval, and generative sketch healing. Our simple design not only yields strong results on all said tasks, but also offers intuitive feature granularity control to tailor for various downstream tasks. Code will be made publicly available. Lan Yang 0014, Kaiyue Pang, Honggang Zhang 0002, Yi-Zhe Song |
ICCV | 1 |
| 2020 | S3Net: Graph Representational Network For Sketch RecognitionabstractSketches are distinctly different to photos. They are highly abstract and exhibit a severe lack of visual cues. Prior works have therefore explored additional traits unique to sketches to help recognition such as stroke ordering. In this paper, we pioneer in studying the role of structure in sketches, for the task of sketch recognition. In particular, we propose a novel graph representation specifically designed for sketches, which follows the inherent hierarchical relationship (segment-stroke-sketch”) of sketching elements. By conforming to this hierarchy, we also introduce ajoint network that encapsulates both the structural and temporal traits of sketches for sketch recognition, termed S3Net.S3Netemploys a recurrent neural network (RNN) to extract segmentlevel features, followed by a graph convolutional network (GCN) to aggregate them into sketch-level features. The RNN first encodes temporal cues in sketches while its outputs are used as node embedding to construct a hierarchical sketch-graph. The GCN module then takes in this sketchgraph to produce a structure-aware embedding for sketches. Extensive experiments on the QuickDraw dataset, exhibit superior performance over state-of-the-arts, surpassing them by over 4%. Ablative studies further demonstrate the effectiveness of the proposed structural graph for both inter-class, and intra-class feature discrimination. Code is available at: https://github.com/yanglan0225/s3net;. Lan Yang 0014, Aneeshan Sain, Linpeng Li, Yonggang Qi, Honggang Zhang 0002, Yi-Zhe Song |
ICME | 1 |
| 2020 | Sketch-SNet: Deeper Subdivision of Temporal Cues for Sketch RecognitionabstractSketch recognition is essential in sketch-related researches. Different from the natural image, the sparse pixel distribution of sketch discards the visual texture which encourages researchers to explore the temporal information of sketch. Using of million-scale datasets, we explore the invariable structure and specific order of strokes in sketch. Prior works based on Recurrent Neural Network (RNN) output different features with changed stroke orders. In particular, we adopt a novel method by employing a Graph Convolutional Network (GCN) to extract invariable structural feature under any orders of strokes. Compared with traditional comprehension of sketch, we further split the temporal information of sketch into two types of feature, invariable structural feature (ISF) and drawing habits feature (DHF) with the aim of finer feature extraction in temporal information. We propose a two-branch GCN-RNN network, Sketch-SNet, to extract two types of feature respectively. The GCN branch is used to extract the ISF through receiving various shuffled strokes of an input sketch. The RNN branch takes the original order to extract DHF by learning the pattern of strokes' order. Extensive experiments on the Quick-Draw dataset demonstrate that our further subdivision of temporal information improves the performance of sketch recognition which surpasses state-of-the-art by a large margin. Yizhou Tan, Lan Yang 0014, Honggang Zhang 0002 |
ICPR | 2 |