Alan Liang

dblp:339/0615 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
0009-0002-9620-8131ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
3D vision · 29% Generative modeling · 16% Reinforcement learning · 16%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 23 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.922026
LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences · AAAI 2026
X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible Controllability · NeurIPS 2025
Computer vision › 3D vision › 3d generation
3d scene generation
1.012026
La La LiDAR: Large-Scale Layout Generation from LiDAR Data · AAAI 2026
Computer vision › 3D vision
3d scene understanding
1.012026
La La LiDAR: Large-Scale Layout Generation from LiDAR Data · AAAI 2026
Computer vision › 3D vision › 3d generation
LiDAR scene generation
1.012026
La La LiDAR: Large-Scale Layout Generation from LiDAR Data · AAAI 2026
Natural language and speech › Language models and text generation
mathematical reasoning
1.012026
TAPO: Dynamic Teacher and Perturbed Answer Injection for Policy Optimization · AAAI 2026
Machine learning › Reinforcement learning
policy optimization
1.012026
TAPO: Dynamic Teacher and Perturbed Answer Injection for Policy Optimization · AAAI 2026
Machine learning › Reinforcement learning
reinforcement learning from human feedback
1.012026
TAPO: Dynamic Teacher and Perturbed Answer Injection for Policy Optimization · AAAI 2026
Machine learning › Reinforcement learning › reward design
reward shaping
1.012026
TAPO: Dynamic Teacher and Perturbed Answer Injection for Policy Optimization · AAAI 2026
Computer vision › Segmentation and scene understanding
scene graph
1.012026
La La LiDAR: Large-Scale Layout Generation from LiDAR Data · AAAI 2026
Computer vision › 3D vision › 3d scene understanding
3d visual grounding
0.912025
3EED: Ground Everything Everywhere in 3D · NeurIPS 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.912025
DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time series · NeurIPS 2025
Machine learning › Generative modeling › diffusion model › controllable generation
controllable scene generation
0.912025
X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible Controllability · NeurIPS 2025
Robotics › Autonomous driving › scenario generation
driving scene generation
0.912025
X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible Controllability · NeurIPS 2025
Computer vision › Vision and language › 3d vision and language
language-guided 3d object localization
0.912025
3EED: Ground Everything Everywhere in 3D · NeurIPS 2025
Machine learning › Representation and self-supervised learning › contrastive learning
multi-view contrastive learning
0.912025
DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time series · NeurIPS 2025
Computer vision › Vision and language
visual grounding
0.912025
3EED: Ground Everything Everywhere in 3D · NeurIPS 2025
Medical and health informatics
clinical time series analysis
0.912025
DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time series · NeurIPS 2025
Robotics › Autonomous driving › perception
LiDAR perception
0.312026
LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences · AAAI 2026
Computer vision › 3D vision › neural rendering
3d gaussian splatting
0.312025
X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible Controllability · NeurIPS 2025
Computer vision › 3D vision
3d scene reconstruction
0.312025
X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible Controllability · NeurIPS 2025
Machine learning › Time series and sequential data
anomaly detection
0.312025
DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time series · NeurIPS 2025
Robotics › Robot navigation and mapping
embodied perception
0.312025
3EED: Ground Everything Everywhere in 3D · NeurIPS 2025
Machine learning › Generative modeling
generative adversarial network
0.312025
DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time series · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

diffusion model · 2.9contrastive learning · 2.7scene graph transformer · 1.0scene graph · 1.0reward shaping · 1.0reinforcement learning · 1.0group relative policy optimization · 1.0autoregressive model · 1.0reconstruction error · 0.9platform-aware normalization · 0.9multi-head attention · 0.9cross-modal alignment · 0.9GAN-enhanced encoder-decoder · 0.9
YearPublicationVenuePosition
2026 TAPO: Dynamic Teacher and Perturbed Answer Injection for Policy Optimization
abstract
Reinforcement learning (RL) has emerged as a powerful framework to improve the reasoning performance of large language models (LLMs), with approaches such as Group Relative Policy Optimization (GRPO) showing promising results. However, GRPO and its variants struggle with collapsed groups (i.e., all-correct or all-incorrect completions), leading to zero-variance rewards and ineffective gradient signals. Moreover, focusing solely on final answer correctness while ignoring the reasoning process, along with rigid length penalties, can hinder training stability and output quality. To address these issues, we introduce TAPO, a reinforcement learning framework that enhances optimization signals by modifying sampled completions within training groups. TAPO incorporates three core techniques: (1) Dynamic Teacher Injection (DTI), which selectively injects high-quality or adversarial examples to restore effective gradient signals in collapsed groups; (2) Perturbed Answer Injection (PAI), which makes partially correct completions to provide contrastive supervision separating reasoning correctness but wrong answer from the trajectories; and (3) InfoLen-Aware Reward Shaping, a fine-grained reward strategy that penalizes outputs based on both length and semantic redundancy, encouraging concise yet informative responses. Extensive experimental results demonstrate that TAPO significantly improves the mathematical reasoning capabilities of LLMs across multiple challenging benchmarks, outperforming the GRPO baseline by a substantial margin. Component-wise ablations further validate the contribution of each proposed technique.
Maowei Jiang, Peter Bús, Moquan Chen, Quangao Liu, Ruiqi Li 0004, Pengyu Zeng, Ruikai Liu, Alan Liang, Yusong Hu, Zhiyong Dong
AAAI11
2026 LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences
abstract
Generative world models have become essential data engines for autonomous driving, yet most focus on videos or occupancy grids and overlook the unique challenges of LiDAR. Extending LiDAR generation to dynamic 4D modeling requires addressing controllability, temporal coherence, and standardized evaluation. We present LiDARCrafter, a unified framework for controllable 4D LiDAR generation and editing. Free-form language instructions are converted into ego-centric scene graphs that guide a tri-branch diffusion model to generate object geometry, motion, and structural priors. An autoregressive module further produces temporally coherent and stable LiDAR sequences with improved global consistency. To enable fair comparison, we introduce a comprehensive benchmark covering scene-, object-, and sequence-level metrics for rigorous and reproducible evaluation. Experiments on nuScenes show that LiDARCrafter achieves state-of-the-art fidelity, controllability, and temporal consistency, paving the way for scalable data augmentation and realistic simulation in diverse scenarios. Code have been publicly available at https://lidarcrafter.github.io.
Alan Liang, Youquan Liu, Dongyue Lu, Lingdong Kong, Huaici Zhao, Wei Tsang Ooi
AAAI1
2026 La La LiDAR: Large-Scale Layout Generation from LiDAR Data
abstract
Controllable generation of realistic LiDAR scenes is crucial for applications such as autonomous driving and robotics. While recent diffusion-based models achieve high-fidelity LiDAR generation, they lack explicit control over foreground objects and spatial relationships, limiting their usefulness for scenario simulation and safety validation. To address these limitations, we propose Large-scale Layout-guided LiDAR generation model ("La La LiDAR"), a novel layout-guided generative framework that introduces semantic-enhanced scene graph diffusion with relation-aware contextual conditioning for structured LiDAR layout generation, followed by foreground-aware control injection for complete scene generation. This enables customizable control over object placement while ensuring spatial and semantic consistency. To support our structured LiDAR generation, we introduce Waymo-SG and nuScenes-SG, two large-scale LiDAR scene graph datasets, along with new evaluation metrics for layout synthesis. Extensive experiments demonstrate that La La LiDAR achieves state-of-the-art performance in both LiDAR generation and downstream perception tasks, establishing a new benchmark for controllable 3D scene generation.
Youquan Liu, Lingdong Kong, Weidong Yang 0001, Xin Li 0110, Alan Liang, Runnan Chen, Ben Fei, Tongliang Liu
AAAI5
2025 A Stacking-Based Ensemble Approach for Predicting Chess Puzzle Difficulty
abstract
FedCSIS 2025 competition is to predict the difficulty of chess puzzles, we present a structured multi-stage regression pipeline developed for the FedCSIS 2025 Challenge.The approach consists of three stages: (i) four Elo-banded base models trained on separate rating ranges to capture localized difficulty semantics and mitigate bias in imbalanced datasets;(ii) a feature-level stacking ensemble combining base predictions with structural attributes, such as success probabilities, failure distributions, and solution length, to enhance cross-band generalization; and (iii) a lightweight post-hoc residual correction to reduce systematic prediction biases.Additionally, an uncertaintyaware mask-based evaluation is introduced to identify the 10% most challenging puzzles for extended scoring.Our method achieved competitive results, ranking 7th in the final leaderboard, while maintaining low computational cost.These findings demonstrate that lightweight, interpretable models, when combined with structural reasoning and uncertainty estimation, can rival more complex deep-learning approaches.This study highlights the potential of structured machine learning pipelines for scalable, human-centric chess puzzle analytics.
Alan Liang, Cenzhi Liu, Ethan Liu
FedCSIS1
2025 Talk2Event: Grounded Understanding of Dynamic Scenes from Event Cameras
abstract
Event cameras offer microsecond-level latency and robustness to motion blur, making them ideal for understanding dynamic environments. Yet, connecting these asynchronous streams to human language remains an open challenge. We introduce Talk2Event, the first large-scale benchmark for language-driven object grounding in event-based perception. Built from real-world driving data, Talk2Event provides over 30,000 validated referring expressions, each enriched with four grounding attributes -- appearance, status, relation to viewer, and relation to other objects -- bridging spatial, temporal, and relational reasoning. To fully exploit these cues, we propose EventRefer, an attribute-aware grounding framework that dynamically fuses multi-attribute representations through a Mixture of Event-Attribute Experts (MoEE). Our method adapts to different modalities and scene dynamics, achieving consistent gains over state-of-the-art baselines in event-only, frame-only, and event-frame fusion settings. We hope our dataset and approach will establish a foundation for advancing multimodal, temporally-aware, and language-driven perception in real-world robotics and autonomy.
Lingdong Kong, Dongyue Lu, Alan Liang, Yuhao Dong, Tianshuai Hu, Lai Xing Ng, Wei Tsang Ooi, Benoit Cottereau
NeurIPS3
2025 3EED: Ground Everything Everywhere in 3D
abstract
Visual grounding in 3D is the key for embodied agents to localize language-referred objects in open-world environments. However, existing benchmarks are limited to indoor focus, single-platform constraints, and small scale. We introduce 3EED, a multi-platform, multi-modal 3D grounding benchmark featuring RGB and LiDAR data from vehicle, drone, and quadruped platforms. We provide over 128,000 objects and 22,000 validated referring expressions across diverse outdoor scenes -- 10x larger than existing datasets. We develop a scalable annotation pipeline combining vision-language model prompting with human verification to ensure high-quality spatial grounding. To support cross-platform learning, we propose platform-aware normalization and cross-modal alignment techniques, and establish benchmark protocols for in-domain and cross-platform evaluations. Our findings reveal significant performance gaps, highlighting the challenges and opportunities of generalizable 3D grounding. The 3EED dataset and benchmark toolkit are released to advance future research in language-driven 3D embodied perception.
Yuhao Dong, Tianshuai Hu, Alan Liang, Youquan Liu, Dongyue Lu, Liang Pan, Lingdong Kong, Junwei Liang 0001, Ziwei Liu 0002
NeurIPS4
2025 DAAC: Discrepancy-Aware Adaptive Contrastive Learning for Medical Time series
abstract
Medical time-series data play a vital role in disease diagnosis but suffer from limited labeled samples and single-center bias, which hinder model generalization and lead to overfitting. To address these challenges, we propose DAAC (Discrepancy-Aware Adaptive Contrastive learning), a learnable multi-view contrastive framework that integrates external normal samples and enhances feature learning through adaptive contrastive strategies. DAAC consists of two key modules: (1) a Discrepancy Estimator, built upon a GAN-enhanced encoder-decoder architecture, captures the distribution of normal data and computes reconstruction errors as indicators of abnormality. These discrepancy features augment the target dataset to mitigate overfitting. (2) an Adaptive Contrastive Learner uses multi-head attention to extract discriminative representations by contrasting embeddings across multiple views and data granularities (subject, trial, epoch, and temporal levels), eliminating the need for handcrafted positive-negative sample pairs. Extensive experiments on three clinical datasets—covering Alzheimer’s disease, Parkinson’s disease, and myocardial infarction—demonstrate that DAAC significantly outperforms existing methods, even when only 10\% of labeled data is available, showing strong generalization and diagnostic performance. Our code is available at https://github.com/CUHKSZ-MED-BioE/DAAC.
Hongfeng Ai, Ruiqi Li 0004, Maowei Jiang, Quangao Liu, Jiahua Dong 0001, Ruiyuan Kang, Alan Liang, Ruikai Liu, Chenzhong Li
NeurIPS8
2025 X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible Controllability
abstract
Diffusion models are advancing autonomous driving by enabling realistic data synthesis, predictive end-to-end planning, and closed-loop simulation, with a primary focus on temporally consistent generation. However, large-scale 3D scene generation requiring spatial coherence remains underexplored. In this paper, we present X-Scene, a novel framework for large-scale driving scene generation that achieves geometric intricacy, appearance fidelity, and flexible controllability. Specifically, X-Scene supports multi-granular control, including low-level layout conditioning driven by user input or text for detailed scene composition, and high-level semantic guidance informed by user intent and LLM-enriched prompts for efficient customization. To enhance geometric and visual fidelity, we introduce a unified pipeline that sequentially generates 3D semantic occupancy and corresponding multi-view images and videos, ensuring alignment and temporal consistency across modalities. We further extend local regions into large-scale scenes via consistency-aware outpainting, which extrapolates occupancy and images from previously generated areas to maintain spatial and visual coherence. The resulting scenes are lifted into high-quality 3DGS representations, supporting diverse applications such as simulation and scene exploration. Extensive experiments demonstrate that X-Scene substantially advances controllability and fidelity in large-scale scene generation, empowering data generation and simulation for autonomous driving.
Yu Yang 0001, Alan Liang, Jianbiao Mei, Yukai Ma, Yong Liu 0007, Gim Hee Lee
NeurIPS2
2024 MyRaft: High Availability in MySQL using Raft
Anirban Rahut, Vinaykumar Bhat, Bartlomiej Pelc, Ahsanul Haque, Yash Botadra, Michael Percy, Ritwik Yadav, Yoshinori Matsunobu, Alan Liang, Igor Pozgaj, Tobias Asplund, Anatoly Karp, Luqun Lou, Pushap Goyal
EDBT13