VLDB 2026 Research / reviewers in the wild / expert
Hanwen Liang
dblp:248/9332
· DBLP profile ↗
14ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PanFlow: Decoupled Motion Control for Panoramic Video GenerationabstractPanoramic video generation has attracted growing attention due to its applications in virtual reality and immersive media. However, existing methods lack explicit motion control and struggle to generate scenes with large and complex motions. We propose PanFlow a novel approach that exploits the spherical nature of panoramas to decouple the highly dynamic camera rotation from the input optical flow condition, enabling more precise control over large and dynamic motions. We further introduce a spherical noise warping strategy to promote loop consistency in motion across panorama boundaries. To support effective training, we curate a large-scale, motion-rich panoramic video dataset with frame-level pose and flow annotations. We also showcase the effectiveness of our method in various applications, including motion transfer and video editing. Extensive experiments demonstrate that PanFlow significantly outperforms prior methods in motion fidelity, visual quality, and temporal coherence. Hanwen Liang, Donny Y. Chen, Qianyi Wu, Konstantinos N. Plataniotis, Camilo Cruz Gambardella, Jianfei Cai 0001 |
AAAI | 2 |
| 2026 | Comp4D: Compositional 4D Scene Generation
Hanwen Liang, Dejia Xu, Neel P. Bhatt, Hezhen Hu, Hanxue Liang, Konstantinos N. Plataniotis |
WACV | 1 |
| 2025 | Wonderland: Navigating 3D Scenes from a Single ImageabstractHow can one efficiently generate high-quality, wide-scope 3D scenes from arbitrary single images? Existing methods suffer several drawbacks, such as requiring multi-view data, time-consuming per-scene optimization, distorted geometry in occluded areas, and low visual quality in backgrounds. Our novel 3D scene reconstruction pipeline overcomes these limitations to tackle the aforesaid challenge. Specifically, we introduce a large-scale reconstruction model that leverages latents from a video diffusion model to predict 3D Gaussian Splattings of scenes in a feed-forward manner. The video diffusion model is designed to create videos precisely following specified camera trajectories, allowing it to generate compressed video latents that encode multi-view information while maintaining 3D consistency. We train the 3D reconstruction model to operate on the video latent space with a progressive learning strategy, enabling the efficient generation of high-quality, wide-scope, and generic 3D scenes. Extensive evaluations across various datasets affirm that our model significantly outperforms existing single-view 3D scene generation methods, especially with out-of-domain images. Thus, we demonstrate for the first time that a 3D reconstruction model can effectively be built upon the latent space of a diffusion model in order to realize efficient 3D scene generation. Project page: https://snap-research.github.io/wonderland/ Hanwen Liang, Junli Cao, Vidit Goel, Guocheng Qian, Sergei Korolev, Demetri Terzopoulos, Konstantinos N. Plataniotis, Sergey Tulyakov, Jian Ren 0005 |
CVPR | 1 |
| 2025 | SLU-DQN: A Model for Anticipatory Steam Detection for Steamer-Filling in Baijiu Intelligent Distillation SystemsabstractThe true implementation of the Anticipatory Steam Detection for Steamer-Filling(ASDSF) process in baijiu intelligent distillation systems, which involves predicting and precisely spreading distillers’ grains before steam emerges, remains a critical unresolved challenge. In this study, we introduce the SLU model, which utilizes SwinLSTM as the core feature extraction module and adopts a U-shaped structure. This model achieves spatiotemporal feature extraction and dynamic change prediction. It is further enhanced by integrating a U-Net module for multi-scale feature fusion and optimized through a Deep Q-Network (DQN)-based decision-making process. The SLU-DQN model, specifically designed for anticipatory material spreading planning in the baijiu Steamer-Filling(SF) distillation system, predicts future steam emission areas. Finally, both quantitative and qualitative experimental results demonstrate the excellent performance of the SLU-DQN model in solving the ASDSF problem. The model achieved 91.1% reward accuracy, an F1-Score of 91% for material spreading point prediction, an MSE of 19.02, and an SSIM of 95.8%. These results not only highlight the model’s superior accuracy in predicting future steam emission areas but also provide a significant technical breakthrough for intelligent baijiu distillation systems, filling a crucial gap in the field. Jiankun Ren, Hanwen Liang, Lizhe Qi, Yunquan Sun |
IROS | 3 |
| 2025 | TiP4GEN: Text to Immersive Panorama 4D Scene GenerationabstractWith the rapid advancement and widespread adoption of VR/AR technologies, there is a growing demand for the creation of high-quality, immersive dynamic scenes. However, existing generation works predominantly concentrate on the creation of static scenes or narrow perspective-view dynamic scenes, falling short of delivering a truly 360-degree immersive experience from any viewpoint. In this paper, we introduce TiP4GEN, an advanced text-to-dynamic panorama scene generation framework that enables fine-grained content control and synthesizes motion-rich, geometry-consistent panoramic 4D scenes. TiP4GEN integrates panorama video generation and dynamic scene reconstruction to create 360-degree immersive virtual environments. For video generation, we introduce a Dual-branch Generation Model consisting of a panorama branch and a perspective branch, responsible for global and local view generation, respectively. A bidirectional cross-attention mechanism facilitates comprehensive information exchange between the branches. For scene reconstruction, we propose a Geometry-aligned Reconstruction Model based on 3D Gaussian Splatting. By aligning spatial-temporal point clouds using metric depth maps and initializing scene cameras with estimated poses, our method ensures geometric consistency and temporal coherence for the reconstructed scenes. Extensive experiments demonstrate the effectiveness of our proposed designs and the superiority of TiP4GEN in generating visually compelling and motion-coherent dynamic panoramic scenes. Hanwen Liang, Dejia Xu, Yuyang Yin, Konstantinos N. Plataniotis, Yao Zhao 0001, Yunchao Wei |
ACM Multimedia | 2 |
| 2025 | Beyond Masked and Unmasked: Discrete Diffusion Models via Partial MaskingabstractMasked diffusion models (MDM) are powerful generative models for discrete data that generate samples by progressively unmasking tokens in a sequence. Each token can take one of two states: masked or unmasked. We observe that token sequences often remain unchanged between consecutive sampling steps; consequently, the model repeatedly processes identical inputs, leading to redundant computation. To address this inefficiency, we propose the Partial masking scheme (Prime), which augments MDM by allowing tokens to take intermediate states interpolated between the masked and unmasked states. This design enables the model to make predictions based on partially observed token information, and facilitates a fine-grained denoising process. We derive a variational training objective and introduce a simple architectural design to accommodate intermediate-state inputs. Our method demonstrates superior performance across a diverse set of generative modeling tasks. On text data, it achieves a perplexity of 15.36 on OpenWebText, outperforming previous MDM (21.52), autoregressive models (17.54), and their hybrid variants (17.58), without relying on an autoregressive formulation. On image data, it attains competitive FID scores of 3.26 on CIFAR-10 and 6.98 on ImageNet-32, comparable to leading continuous generative models. Chen-Hao Chao, Wei-Fang Sun, Hanwen Liang, Chun-Yi Lee, Rahul G. Krishnan |
NeurIPS | 3 |
| 2025 | A Joint Entity and Relation Extraction Model Driven by Fine-Grained Feature ExtractionabstractJoint entity and relation extraction plays a crucial role in various fields, including natural language processing, knowledge graph construction, and question answering systems. Among these methods, the joint extraction approach based on table filling has become a focal point for researchers due to its excellent performance and ability to accurately extract entities and relations from intricate sentences. Despite this, there is still potential to be further explored in such methods. Currently, most methods focus on learning relational features between words but neglect the relational features between labels and between labels and words. These relations are also important because a relation triplet can only be correctly identified when all labels are predicted correctly. This some what increases the risk of missing key information and inefficient decoding in the entity and relation extraction process. To overcome this limitation, we design a fine-grained relation feature extraction module that aims to encourage the model to fully consider the associations between words and labels, as well as between labels, while focusing on the relationships between words. Specifically, we create label representations and embed them into fine-grained relation extraction, enabling it to treat labels and word representations as queries, keys, and values, and then utilize the self-attention mechanism to model the associations between them. Additionally, the table labeling strategy has a significant impact on model performance, especially in terms of decoding efficiency. We propose a vertex labeling method to label the table, enhancing the model’s accuracy in entity recognition and overall decoding efficiency. We evaluate our proposed model on three benchmark datasets. The experimental results demonstrate that our model is effective and achieves state-of-the-art performance on all these datasets. Hongli Yu, Han Cao 0005, Haihang Wang, Hanwen Liang, Yachao Cui |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2024 | AGuided Comparison of Bioinstrumentation Laboratory Data Analysis using Mathematical Software and Generative AIabstractGenerative AI tools are becoming more widely available and have increasing functionality. Students are beginning to integrate AI into their current practice and will need to be able to ethically use AI in their future careers. Instead of banning AI in an undergraduate biomedical instrumentation instructional laboratory course at a large public university, it was intentionally added with awareness of privacy, equity, and accountability. An assignment was adapted to walk students through comparing data analysis by hand, with mathematical software, and with generative AI. The goal of the updated assignment was for students to be able to think critically about the difference between doing analysis by hand, with purpose-built and validated software, and with a generic tool based on a large language model. The submitted post-lab assignments were analyzed by the research team to understand the students' approach to this assignment and what they learned about each method. All the students were able to complete the assignment, however there was mixed feedback on the usefulness of the assignment. Details about the assignment development and analysis of student work on the assignment are included in this paper. Hannah Kimmel, Maya Miriyala, Hanwen Liang, Megha Agrawal, Kaitlyn Tuvilleja, Rebecca M. Reck |
FIE | 3 |
| 2024 | Spot the Difference! Temporal Coarse to Fine to Finer Difference Spotting for Action Recognition in VideosabstractIn this paper, we present a novel difference-spotting strategy for video action recognition inspired by the cognitive challenges posed by the childhood puzzle game "Spot the Difference". Our approach aims to enhance the model’s capability to capture time-series variation and intricate details by gradually integrating distinctive information between action and non-action segments in a temporal "coarse-to-fine-to-finer" manner within a discriminative learning framework. To achieve this, we propose a model-agnostic discriminative learning mechanism that can be easily integrated into existing action recognition networks. Firstly, we incorporate coarse-level discriminative information of action and non-action segments across all videos in a corpus using novel booster nets. Secondly, we introduce a fine-level discrimination objective in the penultimate layer of the network through a novel contrastive learning approach, increasing the distinction between different segments within the same video. Lastly, we incorporate finer discrimination through a novel clip matching mechanism, enhancing the distinction of different consecutive clips within an action segment. Experimental results on multiple benchmark datasets (ActivityNet, HACS, FineAction) and backbone architectures (TSN, TSM, TANet, TPN, Timesformer, VideoSwin) demonstrate the effectiveness of our proposed mechanism. We consistently achieve significant improvements (0.33 - 4%) over the baselines, with competitive single crop results on ActivityNet (87.9%) and HACS (90.21%) datasets. Moreover, our technique achieves stateof-the-art classifier results (94.8%) in the ActivityNet 2022 challenge’s validation set. Yaoxin Li, Deepak Sridhar, Hanwen Liang, Alexander Wong |
ICME | 3 |
| 2024 | An Arc Light Elimination Network Using Polarization and Prior Information
Wenhao Yu 0009, Lizhe Qi, Hanwen Liang, Yunquan Sun |
IJCNN | 3 |
| 2024 | Diffusion4D: Fast Spatial-temporal Consistent 4D generation via Video Diffusion ModelsabstractThe availability of large-scale multimodal datasets and advancements in diffusion models have significantly accelerated progress in 4D content generation. Most prior approaches rely on multiple images or video diffusion models, utilizing score distillation sampling for optimization or generating pseudo novel views for direct supervision. However, these methods are hindered by slow optimization speeds and multi-view inconsistency issues. Spatial and temporal consistency in 4D geometry has been extensively explored respectively in 3D-aware diffusion models and traditional monocular video diffusion models. Building on this foundation, we propose a strategy to migrate the temporal consistency in video diffusion models to the spatial-temporal consistency required for 4D generation. Specifically, we present a novel framework, \textbf{Diffusion4D}, for efficient and scalable 4D content generation. Leveraging a meticulously curated dynamic 3D dataset, we develop a 4D-aware video diffusion model capable of synthesizing orbital views of dynamic 3D assets. To control the dynamic strength of these assets, we introduce a 3D-to-4D motion magnitude metric as guidance. Additionally, we propose a novel motion magnitude reconstruction loss and 3D-aware classifier-free guidance to refine the learning and generation of motion dynamics. After obtaining orbital views of the 4D asset, we perform explicit 4D construction with Gaussian splatting in a coarse-to-fine manner. Extensive experiments demonstrate that our method surpasses prior state-of-the-art techniques in terms of generation efficiency and 4D geometry consistency across various prompt modalities. Hanwen Liang, Yuyang Yin, Dejia Xu, Hanxue Liang, Zhangyang Wang, Konstantinos N. Plataniotis, Yao Zhao 0001, Yunchao Wei |
NeurIPS | 1 |
| 2022 | Self-Supervised Spatiotemporal Representation Learning by Exploiting Video ContinuityabstractRecent self-supervised video representation learning methods have found significant success by exploring essential properties of videos, e.g. speed, temporal order, etc. This work exploits an essential yet under-explored property of videos, the \textit{video continuity}, to obtain supervision signals for self-supervised representation learning. Specifically, we formulate three novel continuity-related pretext tasks, i.e. continuity justification, discontinuity localization, and missing section approximation, that jointly supervise a shared backbone for video representation learning. This self-supervision approach, termed as Continuity Perception Network (CPNet), solves the three tasks altogether and encourages the backbone network to learn local and long-ranged motion and context representations. It outperforms prior arts on multiple downstream tasks, such as action recognition, video retrieval, and action localization. Additionally, the video continuity can be complementary to other coarse-grained video properties for representation learning, and integrating the proposed pretext task to prior arts can yield much performance gains. Hanwen Liang, Niamul Quader, Zhixiang Chi, Lizhe Chen, Peng Dai 0002, Juwei Lu, Yang Wang 0003 |
AAAI | 1 |
| 2021 | Boosting the Generalization Capability in Cross-Domain Few-shot Learning via Noise-enhanced Supervised AutoencoderabstractState of the art (SOTA) few-shot learning (FSL) methods suffer significant performance drop in the presence of domain differences between source and target datasets. The strong discrimination ability on the source dataset does not necessarily translate to high classification accuracy on the target dataset. In this work, we address this cross-domain few-shot learning (CDFSL) problem by boosting the generalization capability of the model. Specifically, we teach the model to capture broader variations of the feature distributions with a novel noise-enhanced supervised autoencoder (NSAE). NSAE trains the model by jointly reconstructing inputs and predicting the labels of inputs as well as their reconstructed pairs. Theoretical analysis based on intra-class correlation (ICC) shows that the feature embeddings learned from NSAE have stronger discrimination and generalization abilities in the target domain. We also take advantage of NSAE structure and propose a two-step fine-tuning procedure that achieves better adaption and improves classification performance in the target domain. Extensive experiments and ablation studies are conducted to demonstrate the effectiveness of the proposed method. Experimental results show that our proposed method consistently outperforms SOTA methods under various conditions. Hanwen Liang, Peng Dai 0002, Juwei Lu |
ICCV | 1 |
| 2020 | New Interpretations of Normalization Methods in Deep LearningabstractIn recent years, a variety of normalization methods have been proposed to help training neural networks, such as batch normalization (BN), layer normalization (LN), weight normalization (WN), group normalization (GN), etc. However, some necessary tools to analyze all these normalization methods are lacking. In this paper, we first propose a lemma to define some necessary tools. Then, we use these tools to make a deep analysis on popular normalization methods and obtain the following conclusions: 1) Most of the normalization methods can be interpreted in a unified framework, namely normalizing pre-activations or weights onto a sphere; 2) Since most of the existing normalization methods are scaling invariant, we can conduct optimization on a sphere with scaling symmetry removed, which can help to stabilize the training of network; 3) We prove that training with these normalization methods can make the norm of weights increase, which could cause adversarial vulnerability as it amplifies the attack. Finally, a series of experiments are conducted to verify these claims. Xiangyong Cao, Hanwen Liang, Weiran Huang 0001, Zewei Chen, Zhenguo Li |
AAAI | 3 |