VLDB 2026 Research / reviewers in the wild / expert
Biao Jiang
dblp:38/4792
· DBLP profile ↗
14ranked-venue papers
7as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorSystems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Generative modeling · 45% 3D vision · 34% Vision and language · 12% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 72% Computer animation and physical simulation · 28% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › diffusion model
human motion generation |
1.4 | 2 | 2024 | MotionChain: Conversational Motion Controllers via Multimodal Prompts · ECCV (26) 2024 Executing your Commands via Motion Diffusion in Latent Space · CVPR 2023 |
Computer vision › 3D vision › 3d shape modeling
3d shape editing |
0.9 | 1 | 2025 | ShapeGPT: 3D Shape Generation With a Unified Multi-Modal Language Model · IEEE Trans. Multim. 2025 |
Computer vision › 3D vision
3d shape modeling |
0.9 | 1 | 2025 | ShapeGPT: 3D Shape Generation With a Unified Multi-Modal Language Model · IEEE Trans. Multim. 2025 |
Computer vision › 3D vision › 3d shape reconstruction
shape completion |
0.9 | 1 | 2025 | ShapeGPT: 3D Shape Generation With a Unified Multi-Modal Language Model · IEEE Trans. Multim. 2025 |
Visual content generation and editing
3d content generation |
0.9 | 1 | 2025 | ShapeGPT: 3D Shape Generation With a Unified Multi-Modal Language Model · IEEE Trans. Multim. 2025 |
Visual content generation and editing › 3d content generation
text-guided 3d shape generation |
0.9 | 1 | 2025 | ShapeGPT: 3D Shape Generation With a Unified Multi-Modal Language Model · IEEE Trans. Multim. 2025 |
Computer vision › Vision and language
multimodal prompt |
0.8 | 1 | 2024 | MotionChain: Conversational Motion Controllers via Multimodal Prompts · ECCV (26) 2024 |
Machine learning › Generative modeling › motion generation
conditional motion generation |
0.7 | 1 | 2023 | Executing your Commands via Motion Diffusion in Latent Space · CVPR 2023 |
Machine learning › Generative modeling
diffusion model |
0.7 | 1 | 2023 | Executing your Commands via Motion Diffusion in Latent Space · CVPR 2023 |
Machine learning › Generative modeling › diffusion model
latent diffusion model |
0.7 | 1 | 2023 | Executing your Commands via Motion Diffusion in Latent Space · CVPR 2023 |
Natural language and speech › Language models and text generation
multimodal language model |
0.7 | 1 | 2023 | MotionGPT: Human Motion as a Foreign Language · NeurIPS 2023 |
Computer animation and physical simulation › motion synthesis
human motion synthesis |
0.7 | 1 | 2023 | MotionGPT: Human Motion as a Foreign Language · NeurIPS 2023 |
Computer vision › Vision and language › motion-language model
motion-language pretraining |
0.2 | 1 | 2023 | MotionGPT: Human Motion as a Foreign Language · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
multimodal alignment · 1.7large language model · 1.7instruction-based generation · 1.7discretization · 1.7vector quantization · 1.3prompt learning · 1.3multimodal prompting · 0.8motion generation · 0.8variational autoencoder · 0.7latent space modeling · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Unified multi-modality conditional latent diffusion model for point cloud generation
Yihang Yang, Zibo Zhao 0001, Fukun Yin, Wen Liu 0003, Yuhan Ding, Biao Jiang, Gang Yu 0002, Tao Chen 0003 |
Pattern Recognit. | 6 |
| 2025 | GLMP: Geometric prior learning with multimodal pre-training representation compensation for 3D human shape estimation
Jilin Huang, Biao Jiang, Mengxi Jiang |
Knowl. Based Syst. | 3 |
| 2025 | ShapeGPT: 3D Shape Generation With a Unified Multi-Modal Language ModelabstractThe advent of large language models, which enable flexibility through instruction-driven approaches, has revolutionized many traditional generative tasks, but large models for 3D data, particularly in comprehensively handling 3D shapes with other modalities, are still under-explored. By achieving instruction-based shape generation, versatile multi-modal generative shape models can significantly benefit various fields, such as 3D virtual construction and network-aided design. In this article, we present ShapeGPT, a shape-included multi-modal framework to leverage strong pre-trained language models to address multiple shape-relevant tasks. Specifically, ShapeGPT employs a “word-sentence-paragraph” framework to discretize continuous shapes into shape words, further assembles these words into shape sentences, and integrates shape with instructional text for multi-modal paragraphs. To learn this shape-language model, we use a three-stage training scheme, including shape representation, multi-modal alignment, and instruction-based generation, to align shape-language codebooks and learn the intricate correlations among these modalities. Extensive experiments demonstrate that ShapeGPT achieves comparable performance across shape-relevant tasks, including text-to-shape, shape-to-text, shape completion, and shape editing. Fukun Yin, Xin Chen 0040, Chi Zhang 0007, Biao Jiang, Zibo Zhao 0001, Wen Liu 0003, Gang Yu 0002, Tao Chen 0003 |
IEEE Trans. Multim. | 4 |
| 2024 | MotionChain: Conversational Motion Controllers via Multimodal Prompts
Biao Jiang, Xin Chen 0040, Chi Zhang 0007, Fukun Yin, Zhuoyuan Li 0006, Gang Yu 0002, Jiayuan Fan 0001 |
ECCV (26) | 1 |
| 2024 | WIP: A Two-Generation Model to Support STEM Education in Hispanic-Serving InstitutionsabstractThis innovative practice WIP paper describes a pioneering National Science Foundation-supported research project designed to address the underrepresentation of minority groups in STEM fields by meeting the educational needs of community college student parents. The Holistic Oasis for Parents' Education (HOPE) Program at Hostos Community College centers on a dual-enrollment model, wherein community college parenting students (HOPE Scholars) pursue their academic goals by taking STEM courses during the summer while their children, ranging from kindergarten through fifth grade, are engaged in enriching hands-on STEM activities on the college campus. By aligning the educational pursuits of parents and children, this program supports the academic advancement of adult learners and fosters a positive learning environment for the next generation. The Program's holistic approach empowers HOPE Scholars to accumulate summer credits in STEM courses and nurtures the STEM talent pipeline by inspiring the younger generation. Biao Jiang, Norberto Michel Hernandez Valdes-Portela, JungHang Lee, Sarah L. Hoiland |
FIE | 1 |
| 2023 | Executing your Commands via Motion Diffusion in Latent SpaceabstractWe study a challenging task, conditional human motion generation, which produces plausible human motion sequences according to various conditional inputs, such as action classes or textual descriptors. Since human motions are highly diverse and have a property of quite different distribution from conditional modalities, such as textual descriptors in natural languages, it is hard to learn a probabilistic mapping from the desired conditional modality to the human motion sequences. Besides, the raw motion data from the motion capture system might be redundant in sequences and contain noises; directly modeling the joint distribution over the raw motion sequences and conditional modalities would need a heavy computational over-head and might result in artifacts introduced by the captured noises. To learn a better representation of the various human motion sequences, we first design a powerful Variational AutoEncoder (VAE) and arrive at a representative and low-dimensional latent code for a human motion sequence. Then, instead of using a diffusion model to establish the connections between the raw motion sequences and the conditional inputs, we perform a diffusion process on the motion latent space. Our proposed Motion Latent-based Diffusion model (MLD) could produce vivid motion sequences conforming to the given conditional inputs and substantially reduce the computational overhead in both the training and inference stages. Extensive experiments on various human motion generation tasks demonstrate that our MLD achieves significant improvements over the state-of-the-art methods among extensive human motion generation tasks, with two orders of magnitude faster than previous diffusion models on raw motion sequences. Xin Chen 0040, Biao Jiang, Wen Liu 0003, Tao Chen 0003, Gang Yu 0002 |
CVPR | 2 |
| 2023 | MVCAL: Multi View Clustering for Active Learning
Biao Jiang |
ICONIP (9) | 2 |
| 2023 | MotionGPT: Human Motion as a Foreign LanguageabstractThough the advancement of pre-trained large language models unfolds, the exploration of building a unified model for language and other multimodal data, such as motion, remains challenging and untouched so far. Fortunately, human motion displays a semantic coupling akin to human language, often perceived as a form of body language. By fusing language data with large-scale motion models, motion-language pre-training that can enhance the performance of motion-related tasks becomes feasible. Driven by this insight, we propose MotionGPT, a unified, versatile, and user-friendly motion-language model to handle multiple motion-relevant tasks. Specifically, we employ the discrete vector quantization for human motion and transfer 3D motion into motion tokens, similar to the generation process of word tokens. Building upon this "motion vocabulary", we perform language modeling on both motion and text in a unified manner, treating human motion as a specific language. Moreover, inspired by prompt learning, we pre-train MotionGPT with a mixture of motion-language data and fine-tune it on prompt-based question-and-answer tasks. Extensive experiments demonstrate that MotionGPT achieves state-of-the-art performances on multiple motion tasks including text-driven motion generation, motion captioning, motion prediction, and motion in-between. Biao Jiang, Xin Chen 0040, Wen Liu 0003, Jingyi Yu 0001, Gang Yu 0002, Tao Chen 0003 |
NeurIPS | 1 |
| 2022 | AEBSR: Active-Sampling and Energy-Based Single Image Super-ResolutionabstractTraditionally, single image super-resolution (SISR) methods randomly crop fixed-size patches in both low and high resolution (LR and HR) images as training samples, and obtain reconstruction model through the regression of LR-HR pixels pairs. However, these will lead to two problems. One is the negligence of the essential information of textures and edges leading to redundant and inefficient training. The other is the lack of the overall perception of data distribution causing poor generalization. To mitigate these issues, we propose Active Sampling and Energy-Based Single Image Super-Resolution (AEBSR), which introduces Active Sampling (AS) and Energy-Based Training (EBT) into SISR. Specifically, we first actively sample texture and edge patches through information entropy, and then align different data distributions between SR and HR images through free energy to perceive the overall distribution characteristics. Extensive experiments show that AS and EBT can further improve the SISR effect and our AEBSR also achieves competitive results compared to the current state-of-the-art SISR approaches. Biao Jiang, Kun Long |
ICIP | 1 |
| 2022 | Unsupervised Unpaired Super-Resolution Using an Active Sampling Strategy Based on Edge DetectionabstractMost existing super-resolution (SR) methods rely on pairs of low resolution (LR) and high resolution (HR) images and predetermined degradation operations (e.g., bicubic), usually trained by supervised learning. However, they often fail in real-world scenarios because of the occurrence of noise and blur. The key reason is that the degradation process is unknown, and no HR-LR pairs can be obtained directly. To address the above issues, this paper explores the optimization of an unsupervised unpaired SR method inspired by generative models such as Generative Adversarial Networks (GAN) and Cycle-Consistent Adversarial Networks (CycleGAN). We propose an active sampling strategy based on edge detection, and introduce a denoising network to construct a novel unsupervised unpaired SR framework. The active sampling strategy can perform image-level feature alignment by sampling image patches actively, thus optimizing the learning direction of the generator. At the same time, the denoising network can reduce the learning difficulty of the generator by preprocessing real-world LR images. Extensive experiments indicate that our method obtains better performance over other existing solutions to the unsupervised unpaired SR challenge, Kun Long, Biao Jiang |
IJCNN | 2 |
| 2020 | Imaging carotid wall mechanical heterogeneity in ultrasound image sequences using Eulerian video magnificationabstractAtherosclerosis is a cardiovascular disease causing gradual infiltration of fatty streaks in the arterial wall and eventually plaque build-up that may rupture and cause a stroke. This induces changes in local biomechanics in response to systemic pressure variations. Such mechanical heterogeneities in the carotid arteries can be studied using ultrasound imaging methods. However, in 2-D ultrasound imaging, mechanical heterogeneity may result in both in-plane motions as well as out-of-plane motion, the latter causing intensity variations. Here, we evaluate linear Eulerian video magnification (EVM) processing on carotid ultrasound image sequences, and its ability to image local mechanical heterogeneity. In addition, we explore the method on several ultrasound image sequences from carotid wall tissues at different atherosclerotic disease stages ranging from healthy, early and late atherosclerosis, as well as changes with pharmacological treatment. The results show that linear EVM can be used to magnify motions in carotid ultrasound image sequences, and to derive heterogeneity maps that can visualize mechanical aspects of the carotid walls and its composition. Empirical case study on carotid walls indicate that the heterogeneity maps transition from homogenic to heterogenic pattern with progression of the atherosclerotic disease. The findings of this work show that mechanical heterogeneity imaging may be important in assessing atherosclerotic disease progression and potential risk prediction. Biao Jiang, Hazrat Ali, Christer Grönlund |
BIBE | 1 |
| 2019 | Deep Neural Network based Visual Inspection with 3D Metric Measurement of Concrete Defects using Wall-climbing RobotabstractThis paper presents a novel metric inspection robot system using a deep neural network to detect and measure surface flaws (i.e., crack and spalling) on concrete structures performed by a wall-climbing robot. The system consists of four modules: robotics data collection module to obtain RGB-D images and IMU measurement, visual-inertial SLAM module to generate pose coupled key-frames with depth information, InspectionNet module to classify each pixel into three classes (back-ground, crack and spalling), and 3D registration and map fusion module to register the flaw patch into registered 3D model overlaid and highlighted with detected flaws for spatial-contextual visualization. The system enables the metric model of each surface flaw patch with pixel-level accuracy and determines its location in 3D space that is significant for structural health assessment and monitoring. The InspectionNet achieves an average accuracy of 87.64% for crack and spalling inspection. We also demonstrate our InspectionNet is robust to view angle, scale and illumination variation. Finally, we design a metric voxel volume map to highlight the flaw in 3D model and provide location and metric information. Bing Li 0008, Guoyong Yang, Yong Chang, Zhaoming Liu, Biao Jiang, Jizhong Xiao |
IROS | 6 |
| 2013 | Real-time Self-teference Video Quality Measurement Technique
Biao Jiang, Tarek N. Saadawi |
WEBIST | 1 |
| 2003 | Exploring UDDI Registries Using Modified OFDAV Browser
Biao Jiang, Mao Lin Huang |
SEKE | 1 |