Yaxuan Song

dblp:357/1472 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0005-2664-4386ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
1 paper
Computing education · 100%
Software engineering, system software, and programming languages
1 paper
Requirements engineering and software design · 50% Program synthesis and code generation · 50%
Human-computer interaction and pervasive computing
2 papers
Ubiquitous computing and smart environments · 64% Learning and educational technologies · 36%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Ubiquitous computing and smart environments
location-based services
0.912025
SCENIC: A Location-based System to Foster Cognitive Development in Children During Car Rides · UIST 2025
Computing education
computational thinking
0.812024
ChatScratch: An AI-Augmented System Toward Autonomous Visual Programming Learning for Children Aged 6-12 · CHI 2024
Computing education
k-12 education
0.812024
ChatScratch: An AI-Augmented System Toward Autonomous Visual Programming Learning for Children Aged 6-12 · CHI 2024
Computing education
programming education
0.812024
ChatScratch: An AI-Augmented System Toward Autonomous Visual Programming Learning for Children Aged 6-12 · CHI 2024
Requirements engineering and software design › model-driven engineering
code generation from design
0.812024
EGFE: End-to-end Grouping of Fragmented Elements in UI Designs with Multimodal Learning · ICSE 2024
Program synthesis and code generation
interface generation
0.812024
EGFE: End-to-end Grouping of Fragmented Elements in UI Designs with Multimodal Learning · ICSE 2024
Learning and educational technologies › AI in education
AI-assisted learning
0.212024
ChatScratch: An AI-Augmented System Toward Autonomous Visual Programming Learning for Children Aged 6-12 · CHI 2024

Methods — techniques the papers use, named apart from their topics

large language model · 1.5image generation · 1.5transformer · 0.8sequence prediction · 0.8multimodal learning · 0.8
YearPublicationVenuePosition
2026 Gene-DML: Dual-Pathway Multi-Level Discrimination for Gene Expression Prediction from Histopathology Images
abstract
Accurately predicting gene expression from histopathology images offers a scalable and non-invasive approach to molecular profiling, with significant implications for precision medicine and computational pathology. However, existing methods often underutilize the cross-modal representation alignment between histopathology images and gene expression profiles across multiple representational levels, thereby limiting their prediction performance. To address this, we propose Gene-DML, a unified framework that structures latent space through Dual-pathway Multi-Level discrimination to enhance correspondence between morphological and transcriptional modalities. The multi-scale instance-level discrimination pathway aligns hierarchical histopathology representations extracted at local, neighbor, and global levels with gene expression profiles, capturing scale-aware morphological-transcriptional relationships. In parallel, the cross-level instance-group discrimination pathway enforces structural consistency between individual (image/gene) instances and modality-crossed (gene/image, respectively) groups, strengthening the alignment across modalities. By jointly modeling fine-grained and structural-level discrimination, Gene-DML is able to learn robust cross-modal representations, enhancing both predictive accuracy and generalization across diverse biological contexts. Extensive experiments on public spatial transcriptomics datasets demonstrate that Gene-DML achieves state-of-the-art performance in gene expression prediction. The code and processed datasets are available at https://github.com/YXSong000/Gene-DML.
Yaxuan Song, Jianan Fan, Hang Chang, Tom Weidong Cai
WACV1
2026 Cell as Point: One-stage framework for efficient cell tracking
abstract
Conventional multi-stage cell tracking approaches rely heavily on detection or segmentation in each frame as a prerequisite, requiring substantial resources for high-quality segmentation masks and increasing the overall prediction time. To address these limitations, we propose CAP , a novel end-to-end one-stage framework that reimagines cell tracking by treating C ell a s P oint. Unlike traditional methods, CAP eliminates the need for explicit detection or segmentation, instead jointly tracking cells for sequences in one stage by leveraging the inherent correlations among their trajectories. This simplification reduces both labeling requirements and pipeline complexity. However, directly processing the entire sequence in one stage poses challenges related to data imbalance in capturing cell division events and long sequence inference. To solve these challenges, CAP introduces two key innovations: (1) adaptive event-guided (AEG) sampling, which prioritizes cell division events to mitigate the occurrence imbalance of cell events, and (2) the rolling-as-window (RAW) inference strategy, which ensures continuous and stable tracking of newly emerging cells over extended sequences. By removing the dependency on segmentation-based preprocessing while addressing the challenges of imbalanced occurrence of cell events and long-sequence tracking, CAP demonstrates promising cell tracking performance and is 8 to 32 times more efficient than existing methods. The code and model checkpoints are available at https://github.com/YXSong000/CAP .
Yaxuan Song, Jianan Fan, Heng Huang 0001, Tom Weidong Cai
Pattern Recognit.1
2025 SCENIC: A Location-based System to Foster Cognitive Development in Children During Car Rides
Liuqing Chen 0002, Yaxuan Song, Ke Lyu, Shuhong Xiao, Yilang Shen, Lingyun Sun
UIST2
2025 From analogy to innovation: A creative conceptual design approach leveraging large language models
Boheng Wang, Haoyu Zuo, Yaxuan Song, Peter R. N. Childs, Liuqing Chen 0002
Adv. Eng. Informatics4
2025 MindScratch: A Visual Programming Support Tool for Classroom Learning Based on Multimodal Generative AI
abstract
Programming is essential in K-12 education and fosters computational thinking skills. Given the complexity of programming and the advanced skills it requires, previous research has introduced user-friendly tools to support young learners. However, our interviews with six programming educators revealed that current tools often fail to reflect classroom learning objectives, offer flexible guidance, and foster creativity. Therefore, we introduced MindScratch, a multimodal generative AI (GAI)-powered visual programming support tool. MindScratch aims to balance structured classroom activities with free programming creation, supporting students in completing creative programming projects based on teacher-set learning objectives while also providing programming scaffolding. The results indicate that, compared to the baseline, MindScratch more effectively helps students achieve high-quality projects aligned with learning objectives. It also enhances students’ computational thinking and thinking. Overall, we believe that GAI-driven educational tools like MindScratch offer students a focused and engaging learning experience.
Yunnong Chen, Shuhong Xiao, Yaxuan Song, Zejian Li, Lingyun Sun, Liuqing Chen 0002
Int. J. Hum. Comput. Interact.3
2024 ChatScratch: An AI-Augmented System Toward Autonomous Visual Programming Learning for Children Aged 6-12
abstract
As Computational Thinking (CT) continues to permeate younger age groups in K-12 education, established CT platforms such as Scratch face challenges in catering to these younger learners, particularly those in the elementary school (ages 6-12). Through formative investigation with Scratch experts, we uncover three key obstacles to children’s autonomous Scratch learning: artist’s block in project planning, bounded creativity in asset creation, and inadequate coding guidance during implementation. To address these barriers, we introduce ChatScratch, an AI-augmented system to facilitate autonomous programming learning for young children. ChatScratch employs structured interactive storyboards and visual cues to overcome artist’s block, integrates digital drawing and advanced image generation technologies to elevate creativity, and leverages Scratch-specialized Large Language Models (LLMs) for professional coding guidance. Our study shows that, compared to Scratch, ChatScratch efficiently fosters autonomous programming learning, and contributes to the creation of high-quality, personally meaningful Scratch projects for children.
Liuqing Chen 0002, Shuhong Xiao, Yunnong Chen, Yaxuan Song, Lingyun Sun
CHI4
2024 EGFE: End-to-end Grouping of Fragmented Elements in UI Designs with Multimodal Learning
abstract
When translating UI design prototypes to code in industry, automatically generating code from design prototypes can expedite the development of applications and GUI iterations. However, in design prototypes without strict design specifications, UI components may be composed of fragmented elements. Grouping these fragmented elements can greatly improve the readability and maintainability of the generated code. Current methods employ a two-stage strategy that introduces hand-crafted rules to group fragmented elements. Unfortunately, the performance of these methods is not satisfying due to visually overlapped and tiny UI elements. In this study, we propose EGFE, a novel method for automatically End-to-end Grouping Fragmented Elements via UI sequence prediction. To facilitate the UI understanding, we innovatively construct a Transformer encoder to model the relationship between the UI elements with multi-modal representation learning. The evaluation on a dataset of 4606 UI prototypes collected from professional UI designers shows that our method outperforms the state-of-the-art baselines in the precision (by 29.75%), recall (by 31.07%), and F1-score (by 30.39%) at edit distance threshold of 4. In addition, we conduct an empirical study to assess the improvement of the generated front-end code. The results demonstrate the effectiveness of our method on a real software engineering application. Our end-to-end fragmented elements grouping method creates opportunities for improving UI-related software engineering tasks.
Liuqing Chen 0002, Yunnong Chen, Shuhong Xiao, Yaxuan Song, Lingyun Sun, Yankun Zhen, Yanfang Chang
ICSE4