VLDB 2026 Research / reviewers in the wild / expert
Yisheng He
dblp:254/0856
· DBLP profile ↗
10ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MulSMo: Multimodal Stylized Motion Generation by Bidirectional Control FlowabstractGenerating motion sequences conforming to a target style while adhering to the given content prompts requires accommodating both the content and style. In existing methods, the information usually only flows from style to content, which may cause conflict between the style and content, harming the integration. Differently, in this work we build a bidirectional control flow between the style and the content, also adjusting the style towards the content, in which case the style-content collision is alleviated and the dynamics of the style is better preserved in the integration. Moreover, we extend the stylized motion generation from one modality, i.e. the style motion, to multiple modalities including texts and images through contrastive learning, leading to flexible style control on the motion generation. To further boost the performance, we advance the motion diffusion to motion-aligned temporal latent diffusion by developing a novel motion VAE. Extensive experiments demonstrate that our method significantly outperforms previous methods across different datasets, while also enabling multimodal signals control. The code of our method will be made publicly available. Zhe Li 0038, Yisheng He, Weichao Shen, Qi Zuo, Lingteng Qiu, Shenhao Zhu, Zilong Dong, Laurence T. Yang, Chang Xu 0002, Weihao Yuan 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | LaMP: Language-Motion Pretraining for Motion Generation, Retrieval, and CaptioningabstractLanguage plays a vital role in the realm of human motion. Existing methods have largely depended on CLIP text embeddings for motion generation, yet they fall short in effectively aligning language and motion due to CLIP’s pretraining on static image-text pairs. This work introduces LaMP, a novel Language-Motion Pretraining model, which transitions from a language-vision to a more suitable language-motion latent space. It addresses key limitations by generating motion-informative text embeddings, significantly enhancing the relevance and semantics of generated motion sequences. With LaMP, we advance three key tasks: text-to-motion generation, motion-text retrieval, and motion captioning through aligned language-motion representation learning. For generation, LaMP instead of CLIP provides the text condition, and an autoregressive masked prediction is designed to achieve mask modeling without rank collapse in transformers. For retrieval, motion features from LaMP’s motion transformer interact with query tokens to retrieve text features from the text transformer, and vice versa. For captioning, we finetune a large language model with the language-informative motion features to develop a strong motion captioning model. In addition, we introduce the LaMP-BertScore metric to assess the alignment of generated motions with textual descriptions. Extensive experimental results on multiple datasets demonstrate substantial improvements over previous methods across all three tasks. Project page: https://aigc3d.github.io/LaMP Zhe Li 0038, Weihao Yuan 0001, Yisheng He, Lingteng Qiu, Shenhao Zhu, Xiaodong Gu 0004, Weichao Shen, Zilong Dong, Laurence T. Yang |
ICLR | 3 |
| 2025 | CoProSketch: Controllable and Progressive Sketch Generation with Diffusion Model
Ruohao Zhan, Yijin Li, Yisheng He, Yichen Shen 0004, Zilong Dong, Guofeng Zhang 0001 |
ACM Multimedia | 3 |
| 2025 | DeepADR: multimodal prediction of adverse drug reaction frequency by integrating early-stage drug discovery information via Kolmogorov-Arnold networksabstractAdverse drug reactions (ADRs) are a major cause of clinical trial failure and postmarket withdrawal, posing significant risks to public health and impeding drug development. While computational methods offer an alternative to costly preclinical testing, existing models often fail with novel compounds by requiring pre-existing information such as drug-ADR associations or by inadequately integrating diverse data sources. Here, we introduce DeepADR, a multimodal deep learning framework for predicting both the occurrence and frequency of ADRs using early-stage, readily available data. DeepADR integrates chemical structures and biological target profiles with semantic representations of ADR terms derived from a large language model (LLMs). These heterogeneous parameters are fused using a Kolmogorov-Arnold Network (KAN), which enhances the modeling of complex, nonlinear relationships among modalities to improve predictive performance. Our model outperforms existing methods in predicting both ADR occurrence and frequency, demonstrating robust generalization to new chemical entities. DeepADR showed consistently better performance than other models across both classification and regression tasks. By effectively integrating chemical, biological, and semantic datasets, DeepADR provides a powerful, scalable tool for the early-stage safety assessment and candidate prioritization. This framework not only facilitates the prioritization of safer drug candidates but also offers a methodology for predicting the toxicity of other hazardous materials, holding significant promise for advancing public health. Jingting Wan, Chenyang Jia, Danhong Dong, Yigang Chen 0001, Yang-Chi-Dung Lin, Yisheng He, Hsi-Yuan Huang, Hsien-Da Huang |
Briefings Bioinform. | 6 |
| 2024 | Freditor: High-Fidelity and Transferable NeRF Editing by Frequency Decomposition
Yisheng He, Weihao Yuan 0001, Siyu Zhu 0001, Zilong Dong, Liefeng Bo, Qixing Huang |
ECCV (41) | 1 |
| 2024 | MoGenTS: Motion Generation based on Spatial-Temporal Joint ModelingabstractMotion generation from discrete quantization offers many advantages over continuous regression, but at the cost of inevitable approximation errors. Previous methods usually quantize the entire body pose into one code, which not only faces the difficulty in encoding all joints within one vector but also loses the spatial relationship between different joints. Differently, in this work we quantize each individual joint into one vector, which i) simplifies the quantization process as the complexity associated with a single joint is markedly lower than that of the entire pose; ii) maintains a spatial-temporal structure that preserves both the spatial relationships among joints and the temporal movement patterns; iii) yields a 2D token map, which enables the application of various 2D operations widely used in 2D images. Grounded in the 2D motion quantization, we build a spatial-temporal modeling framework, where 2D joint VQVAE, temporal-spatial 2D masking technique, and spatial-temporal 2D attention are proposed to take advantage of spatial-temporal signals among the 2D tokens. Extensive experiments demonstrate that our method significantly outperforms previous methods across different datasets, with a $26.6\%$ decrease of FID on HumanML3D and a $29.9\%$ decrease on KIT-ML. Weihao Yuan 0001, Yisheng He, Weichao Shen, Xiaodong Gu 0004, Zilong Dong, Liefeng Bo, Qixing Huang |
NeurIPS | 2 |
| 2024 | GIC: Gaussian-Informed Continuum for Physical Property Identification and SimulationabstractThis paper studies the problem of estimating physical properties (system identification) through visual observations. To facilitate geometry-aware guidance in physical property estimation, we introduce a novel hybrid framework that leverages 3D Gaussian representation to not only capture explicit shapes but also enable the simulated continuum to render object masks as 2D shape surrogates during training. We propose a new dynamic 3D Gaussian framework based on motion factorization to recover the object as 3D Gaussian point sets across different time states. Furthermore, we develop a coarse-to-fine filling strategy to generate the density fields of the object from the Gaussian reconstruction, allowing for the extraction of object continuums along with their surfaces and the integration of Gaussian attributes into these continuum. In addition to the extracted object surfaces, the Gaussian-informed continuum also enables the rendering of object masks during simulations, serving as 2D-shape guidance for physical property estimation. Extensive experimental evaluations demonstrate that our pipeline achieves state-of-the-art performance across multiple benchmarks and metrics. Additionally, we illustrate the effectiveness of the proposed method through real-world demonstrations, showcasing its practical utility. Our project page is at https://jukgei.github.io/project/gic. Junhao Cai, Yuji Yang, Weihao Yuan 0001, Yisheng He, Zilong Dong, Liefeng Bo, Qifeng Chen 0001 |
NeurIPS | 4 |
| 2022 | FS6D: Few-Shot 6D Pose Estimation of Novel Objectsabstract6D object pose estimation networks are limited in their capability to scale to large numbers of object instances due to the close-set assumption and their reliance on high-fidelity object CAD models. In this work, we study a new open set problem; the few-shot 6D object poses estimation: estimating the 6D pose of an unknown object by a few support views without extra training. To tackle the problem, we point out the importance of fully exploring the appearance and geometric relationship between the given support views and query scene patches and propose a dense prototypes matching framework by extracting and matching dense RGBD prototypes with transformers. Moreover, we show that the priors from diverse appearances and shapes are crucial to the generalization capability under the problem setting and thus propose a large-scale RGBD photorealistic dataset (ShapeNet6D) for network pre-training. A simple and effective online texture blending approach is also introduced to eliminate the domain gap from the synthesis dataset, which enriches appearance diversity at a low cost. Finally, we discuss possible solutions to this problem and establish benchmarks on popular datasets to facilitate future research. [project page] Yisheng He, Haoqiang Fan, Jian Sun 0001, Qifeng Chen 0001 |
CVPR | 1 |
| 2021 | FFB6D: A Full Flow Bidirectional Fusion Network for 6D Pose EstimationabstractIn this work, we present FFB6D, a Full Flow Bidirectional fusion network designed for 6D pose estimation from a single RGBD image. Our key insight is that appearance information in the RGB image and geometry information from the depth image are two complementary data sources, and it still remains unknown how to fully leverage them. Towards this end, we propose FFB6D, which learns to combine appearance and geometry information for representation learning as well as output representation selection. Specifically, at the representation learning stage, we build bidirectional fusion modules in the full flow of the two networks, where fusion is applied to each encoding and decoding layer. In this way, the two networks can leverage local and global complementary in-formation from the other one to obtain better representations. Moreover, at the output representation stage, we designed a simple but effective 3D keypoints selection algorithm considering the texture and geometry information of objects, which simplifies keypoint localization for precise pose estimation. Experimental results show that our method outperforms the state-of-the-art by large margins on several benchmarks. Code and video are available at https://github.com/ethnhe/FFB6D.git. Yisheng He, Haoqiang Fan, Qifeng Chen 0001, Jian Sun 0001 |
CVPR | 1 |
| 2020 | PVN3D: A Deep Point-Wise 3D Keypoints Voting Network for 6DoF Pose EstimationabstractIn this work, we present a novel data-driven method for robust 6DoF object pose estimation from a single RGBD image. Unlike previous methods that directly regressing pose parameters, we tackle this challenging task with a keypoint-based approach. Specifically, we propose a deep Hough voting network to detect 3D keypoints of objects and then estimate the 6D pose parameters within a least-squares fitting manner. Our method is a natural extension of 2D-keypoint approaches that successfully work on RGB based 6DoF estimation. It allows us to fully utilize the geometric constraint of rigid objects with the extra depth information and is easy for a network to learn and optimize. Extensive experiments were conducted to demonstrate the effectiveness of 3D-keypoint detection in the 6D pose estimation task. Experimental results also show our method outperforms the state-of-the-art methods by large margins on several benchmarks. Code and video are available at https://github.com/ethnhe/PVN3D.git. Yisheng He, Wei Sun 0029, Jianran Liu, Haoqiang Fan, Jian Sun 0001 |
CVPR | 1 |