EDBT 2026 Demo / reviewers in the wild / expert
Dongliang Kou
dblp:348/9444
· DBLP profile ↗
10ranked-venue papers
0as first author
10since 2021 · last 2026
0009-0009-5792-6895ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniMGS: Unifying Mesh and 3D Gaussian Splatting with Single-Pass Rasterization and Proxy-Based DeformationabstractJoint rendering and deformation of mesh and 3D Gaussian Splatting (3DGS) have significant value as both representations offer complementary advantages for graphics applications. However, due to differences in representation and rendering pipelines, existing studies render meshes and 3DGS separately, making it difficult to accurately handle occlusions and transparency. Moreover, the deformed 3DGS still suffers from visual artifacts due to the sensitivity to the topology quality of the proxy mesh. These issues pose serious obstacles to the joint use of 3DGS and meshes, making it difficult to adapt 3DGS to conventional mesh-oriented graphics pipelines. We propose UniMGS, the first unified framework for rasterizing mesh and 3DGS in a single-pass anti-aliased manner, with a novel binding strategy for 3DGS deformation based on proxy mesh. Our key insight is to blend the colors of both triangle and Gaussian fragments by anti-aliased α-blending in a single pass, achieving visually coherent results with precise handling of occlusion and transparency. To improve the visual appearance of the deformed 3DGS, our Gaussian-centric binding strategy employs a proxy mesh and spatially associates Gaussians with the mesh faces, significantly reducing rendering artifacts. With these two components, UniMGS enables the visualization and manipulation of 3D objects represented by mesh or 3DGS within a unified framework, opening up new possibilities in embodied AI, virtual reality, and gaming. We will release our source code to facilitate future research. Zeyu Xiao 0001, Yimin Cong, Dongliang Kou, Zhenyi Wu, Dingkang Yang, Peng Zhai, Lihua Zhang 0002 |
AAAI | 5 |
| 2025 | MAFD: Fine-Grained Motion Style Transfer with Adaptive Signal FusionabstractMotion style transfer allows for the swift switching of different styles within the same motion for virtual avatars, offering significant efficiency gains and enhanced motion diversity compared to traditional motion capture methods. However, many existing methods struggle with controlling fine details in complex motions, leading to models that capture only coarse-grained style characteristics. To overcome this limitation, we introduce the Motion Adaptive Fusion Diffusion (MAFD) framework, which leverages adaptive signal fusion to highlight essential style-defining features while minimizing redundant information. Moreover, current diffusion-based denoisers often fail to effectively capture the temporal relationships in motion sequences, producing rigid and fragmented stylized motions. Drawing inspiration from the Mamba model, we propose the Style Mamba Denoiser (SMD), which adopts a selection mechanism to preserve long-range dependencies and maintain temporal coherence. Extensive experiments show that our approach outperforms state-of-the-art methods in both qualitative and quantitative evaluations, achieving more refined and coherent stylized motions. Ziyun Qian, Dingkang Yang, Mingcheng Li, Dongliang Kou, Lihua Zhang 0002 |
ICASSP | 4 |
| 2025 | SAF: Local Shape-aware Face-based Garment Collision Handling via Neural SDFsabstractLearning-based garment prediction presents an appealing alternative to physics-based methods owing to its high efficiency. However, the predicted garments can exhibit noticeable penetrations into the body. Many collision-handling methods operate in the point domain, which has an inherent limitation in addressing the penetration issues with the mesh faces. In light of this, we propose a local Shape-Aware, Face-based collision handling approach (SAF) that can be applied to garment prediction networks to achieve real-time collision handling. Considering the nature of garments, we design a compact formulation to model the continuity of the garment surface, and utilize neural Signed Distance Fields (SDFs) to accomplish penetration resolution. Recognizing the significant impact of the body’s local shape on collision handling, we further propose using the angles between SDF gradients to characterize the sharpness of the body. Our approach can compensate for the inaccuracy of neural SDFs and preserve local smoothness and details. Extensive experiments demonstrate the outstanding performance and generalizability of our method. Minzhe Tang, Ruisheng Yuan, Dongliang Kou, Lihua Zhang 0002 |
ICASSP | 3 |
| 2025 | Robust Signed Distance Fields for Articulated Human Body Reconstruction via Multiresolution Hash EncodingabstractThe Signed Distance Fields (SDF) of the human body has broad applications in shape representation, collision handling, and medical image analysis, etc. However, due to the inherently high complexity of human motion, computing the SDFs of dynamic human bodies both accurately and efficiently has long been a challenging problem in computer graphics. In this paper, we demonstrate that by decomposing a widely used explicit human body model (SMPL) and modeling each component in a targeted manner, we can simultaneously realize both efficiency and accuracy. From a high level, the pipeline of the explicit model can be divided into Linear Blend Skinning (LBS) and Pose Space Deformation (PSD). By partitioning the human body into multiple parts and using the transformation matrix of each part, we apply inverse transformations to map spatial points from the posed space back to the canonical space. This eliminates the need for learning transformations and significantly reduces the difficulty of learning PSD. We observe that PSD is essentially a weighted sum of a series of fixed corrective shapes, where the only variable is the coefficient. We propose using Multiresolution Hash Encoding (MHE) to accurately capture the influence of each corrective shape on the SDF and aggregate the features in a manner similar to the explicit model. Our experiments show that our method is robust, effective, and highly efficient. Minzhe Tang, Dongliang Kou, Mingcheng Li, Lihua Zhang 0002 |
IJCNN | 2 |
| 2025 | UMSD: High Realism Motion Style Transfer via Unified Mamba-based DiffusionabstractMotion style transfer is a significant research area in computer vision, enabling the rapid switching of stylistic variations for the same motion in virtual digital humans. This dramatically enhances the richness and realism of motions, making it widely applicable in multimedia contexts such as film, gaming, and the Metaverse. However, most existing methods employ a two-stream structure, which often overlooks the intrinsic relationships between content and style motions, resulting in information loss and misalignment. Additionally, these methods struggle to capture temporal dependencies in long-range motion sequences, resulting in less natural outputs. To address these limitations, we propose a Unified Motion Style Diffusion (UMSD) Framework that simultaneously extracts features from content and style motions, achieving comprehensive information interaction. We also introduce the Motion Style Mamba (MSM) denoiser, which, for the first time in motion style transfer, leverages Mamba's powerful sequence modelling capability to produce more temporally coherent stylized motion sequences. Furthermore, we design a diffusion-based content consistency loss and a style consistency loss to ensure that the framework preserves content motion while effectively learning style motion features. Extensive experiments demonstrate that our approach outperforms State-Of-The-Art (SOTA) methods qualitatively and quantitatively, achieving more realistic and coherent motion style transfer. Ziyun Qian, Zeyu Xiao 0001, Xingliang Jin, Dingkang Yang, Mingcheng Li, Zhenyi Wu, Dongliang Kou, Peng Zhai, Lihua Zhang 0002 |
ACM Multimedia | 7 |
| 2024 | Correlation-Decoupled Knowledge Distillation for Multimodal Sentiment Analysis with Incomplete ModalitiesabstractMultimodal sentiment analysis (MSA) aims to understand human sentiment through multimodal data. Most MSA efforts are based on the assumption of modality completeness. However, in real-world applications, some practical factors cause uncertain modality missingness, which drastically degrades the model's performance. To this end, we propose a Correlation-decoupled Knowledge Distillation (CorrKD) framework for the MSA task under uncertain missing modalities. Specifically, we present a sample-level contrastive distillation mechanism that transfers comprehensive knowledge containing cross-sample correlations to reconstruct missing semantics. Moreover, a category-guided prototype distillation mechanism is introduced to capture cross-category correlations using category prototypes to align feature distributions and generate favorable joint representations. Eventually, we design a response-disentangled consistency distillation strategy to optimize the sentiment decision boundaries of the student network through response disentanglement and mutual information maximization. Comprehensive experiments on three datasets indicate that our framework can achieve favorable improvements compared with several baselines. Mingcheng Li, Dingkang Yang, Shuaibing Wang, Yan Wang 0068, Kun Yang 0010, Dongliang Kou, Ziyun Qian, Lihua Zhang 0002 |
CVPR | 8 |
| 2024 | IIPC: Intra-Inter Patch Correlations for Garment Collision HandlingabstractRealistic garment simulation is critical for digital humans. However, noticeable penetrations still exist in current learning-based garment simulation techniques. To reduce penetrations in predicted garments, we resort to the garment geometry and neural Signed Distance Fields (SDFs) for effective collision handling. The key idea of our method is that we divide the garment into patches and model the local and global garment geometry through Intra- and Inter-Patch Correlations (IIPC), which can be easily learned through the powerful context-understanding ability of Transformers. The geometry information is then utilized to predict a per-vertex moving offset, according to which we move the penetrating vertices along the SDF’s gradient directions to solve collisions. Our module can be coupled with learning-based backbones to effectively solve penetrations while retaining real-time performance. Extensive experiments show that the proposed method excels the prior works significantly. Ruisheng Yuan, Minzhe Tang, Dongliang Kou, Dingkang Yang, Lihua Zhang 0002 |
ICME | 3 |
| 2024 | IF-Garments: Reconstructing Your Intersection-Free Multi-Layered Garments from Monocular VideosabstractReconstructing garments from monocular videos has attracted considerable attention as it provides a convenient and low-cost solution for clothing digitization. In reality, people wear clothing with countless variations and multiple layers. Existing studies attempt to extract garments from a single video. They either behave poorly in generalization due to reliance on limited clothing templates or struggle to handle the intersections of multi-layered clothing leading to the lack of physical plausibility. Besides, there are inevitable and undetectable overlaps for a single video that hinder researchers from modeling complete and intersection-free multi-layered clothing. To address the above limitations, in this paper, we propose a novel method to reconstruct multi-layered clothing from multiple monocular videos sequentially, which surpasses existing work in generalization and robustness against penetration. For each video, neural fields are employed to implicitly represent the clothed body, from which the meshes with frame-consistent structures are explicitly extracted. Next, we implement a template-free method for extracting a single garment by back-projecting the image segmentation labels of different frames onto these meshes. In this way, multiple garments can be obtained from these monocular videos and then aligned to form the whole outfit. However, intersection always occurs due to overlapping deformation in the real world and perceptual errors in monocular videos. To this end, we innovatively introduce a physics-aware module that combines neural fields with a position-based simulation framework to fine-tune the penetrating vertices of garments, ensuring robustly intersection-free. Additionally, we collect a mini dataset with fashionable garments to evaluate the quality of clothing reconstruction comprehensively. We release our code and data at https://github.com/SMY19999/IF-Garments. Qipeng Yan, Zhuoer Liang, Dongliang Kou, Dingkang Yang, Ruisheng Yuan, Mingcheng Li, Lihua Zhang 0002 |
ACM Multimedia | 4 |
| 2024 | Toward Robust Incomplete Multimodal Sentiment Analysis via Hierarchical Representation LearningabstractMultimodal Sentiment Analysis (MSA) is an important research area that aims to understand and recognize human sentiment through multiple modalities. The complementary information provided by multimodal fusion promotes better sentiment analysis compared to utilizing only a single modality. Nevertheless, in real-world applications, many unavoidable factors may lead to situations of uncertain modality missing, thus hindering the effectiveness of multimodal modeling and degrading the model’s performance. To this end, we propose a Hierarchical Representation Learning Framework (HRLF) for the MSA task under uncertain missing modalities. Specifically, we propose a fine-grained representation factorization module that sufficiently extracts valuable sentiment information by factorizing modality into sentiment-relevant and modality-specific representations through crossmodal translation and sentiment semantic reconstruction. Moreover, a hierarchical mutual information maximization mechanism is introduced to incrementally maximize the mutual information between multi-scale representations to align and reconstruct the high-level semantics in the representations. Ultimately, we propose a hierarchical adversarial learning mechanism that further aligns and adapts the latent distribution of sentiment-relevant representations to produce robust joint multimodal representations. Comprehensive experiments on three datasets demonstrate that HRLF significantly improves MSA performance under uncertain modality missing cases. Mingcheng Li, Dingkang Yang, Yang Liu 0246, Shunli Wang 0001, Jiawei Chen 0012, Shuaibing Wang, Jinjie Wei, Qingyao Xu, Xiaolu Hou, Ziyun Qian, Dongliang Kou, Lihua Zhang 0002 |
NeurIPS | 13 |
| 2024 | Multimodal Token Fusion and Optimization for 3D Human Mesh Reconstruction with Transformers
Sunli Wang, Dongliang Kou, Qiangbin Xie, Lihuang Zhang |
PRCV (6) | 4 |