Kai Wang 0012

dblp:78/2022-12 · DBLP profile ↗
← Back
28ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0002-1171-0281ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 8 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment
abstract
Contrastive vision-language models like CLIP have achieved impressive results in image-text retrieval by aligning image and text representations in a shared embedding space. However, these models often treat text as flat sequences, limiting their ability to handle complex, compositional, and long-form descriptions. In particular, they fail to capture two essential properties of language: semantic hierarchy, which reflects the multi-level compositional structure of text, and semantic monotonicity, where richer descriptions should result in stronger alignment with visual content. To address these limitations, we propose HiMo-CLIP, a representation-level framework that enhances CLIP-style models without modifying the encoder architecture. HiMo-CLIP introduces two key components: a hierarchical decomposition (HiDe) module that extracts latent semantic components from long-form text via in-batch PCA, enabling flexible, batch-aware alignment across different semantic granularities, and a monotonicity-aware contrastive loss (MoLo) that jointly aligns global and component-level representations, encouraging the model to internalize semantic ordering and alignment strength as a function of textual completeness. These components work together to produce structured, cognitively aligned cross-modal representations. Experiments on multiple image-text retrieval benchmarks show that HiMo-CLIP consistently outperforms strong baselines, particularly under long or compositional descriptions.
Ruijia Wu, Fei Shen 0004, Shaoan Zhao, Qiang Hui, Huanlin Gao, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian
AAAI10
2026 Enhanced data techniques and optimization in conversational gesture generation
Xiang Wang 0018, Yifeng Peng, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian
CCF Trans. Pervasive Comput. Interact.4
2026 KAConvNet: Kolmogorov-Arnold convolutional networks for vision recognition
Zhaoxiang Liu, Zhicheng Ma, Kaikai Zhao, Kai Wang 0012, Shiguo Lian
Image Vis. Comput.4
2025 Optimizing for the Shortest Path in Denoising Diffusion Model
abstract
In this research, we propose a novel denoising diffusion model based on shortest-path modeling that optimizes residual propagation to enhance both denoising efficiency and quality. Drawing on Denoising Diffusion Implicit Models (DDIM) and insights from graph theory, our model, termed the Shortest Path Diffusion Model (ShortDF), treats the denoising process as a shortest-path problem aimed at minimizing reconstruction error. By optimizing the initial residuals, we improve the efficiency of the reverse diffusion process and the quality of the generated samples. Extensive experiments on multiple standard benchmarks demonstrate that ShortDF significantly reduces diffusion time (or steps) while enhancing the visual fidelity of generated samples compared to prior arts. This work, we suppose, paves the way for interactive diffusion-based applications and establishes a foundation for rapid data generation. Code is available at https://github.com/UnicomAI/ShortDF.
Xingpeng Zhang, Zhaoxiang Liu, Kai Wang 0012, Min Wang 0031, Yanlin Qian, Shiguo Lian
CVPR6
2025 ILearnRobot: An Interactive Learning-Based Multi-modal Robot with Continuous Improvement
Kohou Wang, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian
ICIC (14)7
2025 Art3D-Fusion: A Hybrid Framework for Visual Synthesis with Artistic Control
Kohou Wang, Zhaoxiang Liu, Zezhou Chen, Xin Wang 0135, Kai Wang 0012, Shiguo Lian
ICIG (1)8
2025 Data Leakage Detection in Large Vision-Language Models via Multimodal Perturbation
Xin Wang 0135, Zhaoxiang Liu, Yue Zhan, Kaikai Zhao, Kai Wang 0012, Shiguo Lian
ICIG (1)5
2025 AV-DiT: Taming Image Diffusion Transformers for Efficient Joint Audio and Video Generation
abstract
Recent Diffusion Transformers (DiTs) have shown impressive capabilities in generating single-modality content, including images, videos, and audio. However, the potential of DiTs to enable superb multimodal content creation remains underexplored. To bridge this gap, we introduce AV-DiT, a novel and efficient audio-visual diffusion transformer designed to generate high-quality, realistic videos with synchronized audio tracks. To minimize model complexity and computational costs, our AV-DiT utilizes a modality-shared DiT backbone pre-trained on image-only data, with only newly inserted adapters being trainable. This shared backbone facilitates the generation of both audio and video. Specifically, the video branch incorporates a trainable temporal attention layer into a pre-trained DiT block for capturing the temporal consistency for video generation. In addition, a small number of trainable parameters adapt the image-based DiT block to learn the acoustic characteristics for audio generation. An extra shared self-attention block reused from the DiT block, equipped with lightweight parameters, facilitates feature interaction between audio and visual modalities for alignment. Extensive experiments on the datasets demonstrate that our AV-DiT achieves state-of-the-art performance in joint audio-visual generation with significantly fewer tunable parameters. Furthermore, our results highlight that a single shared image generative backbone with modality-specific adaptations is sufficient for constructing a joint audio-video generator.
Kai Wang 0012, Shijian Deng, Jing Shi 0005, Dimitrios Hatzinakos, Yapeng Tian
ACM Multimedia1
2025 CP3: Customizable 3D Pop-Out Effect Creation for Immersive Content Using Multimodal Models
abstract
In this paper, a multi-modal model based 3D pop-out video generation framework (CP3) is proposed to solve the shortcomings of the existing video generation technology for accurate control of 3D pop-out effects. 3D pop-out effects create an immersive visual experience by changing the disparity of a particular object so that it appears beyond the screen. However, although software has made some progress in this area, there is currently no effective way to accurately control 3D pop-out effects and generate high-quality video. In addition, the lack of high-quality 3D pop-out effect data sets is also one of the bottlenecks in the field. Therefore, the CP3 framework proposed in this paper utilizes multi-modal models to help 3D video creators make 3D pop-out effects, enhance the audience's sense of immersion and visual comfort, and thus promote the development of 3D effect generation technology. To support the training and evaluation of this framework, a new dataset containing 37000 frames of pop-out effects is constructed, such as text guidance, segmentation results, depth maps, optical flow, and the trajectory of the pop-out target. Through the 3D UNet model based on the potential de-noising diffusion mechanism, combined with the 3D-try module in the CP3 framework and Mask Encoder, this paper has achieved remarkable results in the generation of 3D pop-out effect videos. The results of the experiment show that the CP3 framework demonstrates its advantages in generating immersive 3D pop-out effects in comparison to existing technologies.
Zezhou Chen, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian
ACM Multimedia7
2025 LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation
abstract
We present LeMiCa, a training-free and efficient acceleration framework for diffusion-based video generation. While existing caching strategies primarily focus on reducing local heuristic errors, they often overlook the accumulation of global errors, leading to noticeable content degradation between accelerated and original videos. To address this issue, we formulate cache scheduling as a directed graph with error-weighted edges and introduce a Lexicographic Minimax Path Optimization strategy that explicitly bounds the worst-case path error. This approach substantially improves the consistency of global content and style across generated frames. Extensive experiments on multiple text-to-video benchmarks demonstrate that LeMiCa delivers dual improvements in both inference speed and generation quality. Notably, our method achieves a 2.9× speedup on the Latte model and reaches an LPIPS score of 0.05 on Open-Sora, outperforming prior caching techniques. Importantly, these gains come with minimal perceptual quality degradation, making LeMiCa a robust and generalizable paradigm for accelerating diffusion-based video generation. We believe this approach can serve as a strong foundation for future research on efficient and reliable video synthesis.
Huanlin Gao, Fuyuan Shi, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian
NeurIPS7
2025 Joint Deblurring and 3D Reconstruction for Macrophotography
abstract
Abstract Macro lens has the advantages of high resolution and large magnification, and 3D modeling of small and detailed objects can provide richer information. However, defocus blur in macrophotography is a long‐standing problem that heavily hinders the clear imaging of the captured objects and high‐quality 3D reconstruction of them. Traditional image deblurring methods require a large number of images and annotations, and there is currently no multi‐view 3D reconstruction method for macrophotography. In this work, we propose a joint deblurring and 3D reconstruction method for macrophotography. Starting from multi‐view blurry images captured, we jointly optimize the clear 3D model of the object and the defocus blur kernel of each pixel. The entire framework adopts a differentiable rendering method to self‐supervise the optimization of the 3D model and the defocus blur kernel. Extensive experiments show that from a small number of multi‐view images, our proposed method can not only achieve high‐quality image deblurring but also recover high‐fidelity 3D appearance.
Liangchen Li, Yuqi Zhou 0004, Kai Wang 0012, Juyong Zhang
Comput. Graph. Forum4
2025 PSTF-AttControl: Per-subject-tuning-free personalized image generation with controllable face attributes
Zhaoxiang Liu, Zezhou Chen, Kai Wang 0012, Shiguo Lian
Image Vis. Comput.7
2025 MITS: A large-scale multimodal benchmark dataset for Intelligent Traffic Surveillance
Kaikai Zhao, Zhaoxiang Liu, Xin Wang 0135, Zhicheng Ma, Yajun Xu, Wenjing Zhang 0006, Yibing Nan, Kai Wang 0012, Shiguo Lian
Image Vis. Comput.9
2024 A Multimodal Benchmark Dataset and Model for Crop Disease Diagnosis
Zhaoxiang Liu, Zezhou Chen, Kohou Wang, Kai Wang 0012, Shiguo Lian
ECCV (86)6
2024 Spatial-Temporal Transformer Network for Continuous Action Recognition in Industrial Assembly
Shanghua Tang, Shaoan Zhao, Yimin Lin, Kai Wang 0012, Zhaoxiang Liu, Shiguo Lian
ICIC (10)8
2024 Self-supervised Visual Anomaly Detection with Image Patch Generation and Comparison Networks
Kaikai Zhao, Yimin Lin, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian
ICIC (10)6
2024 A Large Vision-Language Model based Environment Perception System for Visually Impaired People
abstract
It is a challenging task for visually impaired people to perceive their surrounding environment due to the complexity of the natural scenes. Their personal and social activities are thus highly limited. This paper introduces a Large Vision-Language Model(LVLM) based environment perception system which helps them to better understand the surrounding environment, by capturing the current scene they face with a wearable device, and then letting them retrieve the analysis results through the device. The visually impaired people could acquire a global description of the scene by long pressing the screen to activate the LVLM output, retrieve the categories of the objects in the scene resulting from a segmentation model by tapping or swiping the screen, and get a detailed description of the objects they are interested in by double-tapping the screen. To help visually impaired people more accurately perceive the world, this paper proposes incorporating the segmentation result of the RGB image as external knowledge into the input of LVLM to reduce the LVLM’s hallucination. Technical experiments on POPE, MME and LLaVA-QA90 show that the system could provide a more accurate description of the scene compared to Qwen-VL-Chat, exploratory experiments show that the system helps visually impaired people to perceive the surrounding environment effectively.
Zezhou Chen, Zhaoxiang Liu, Kai Wang 0012, Kohou Wang, Shiguo Lian
IROS3
2024 Reparameterization-Based Parameter-Efficient Fine-Tuning Methods for Large Language Models: A Systematic Survey
Zezhou Chen, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian
NLPCC (3)3
2024 What is the Best Model? Application-Driven Evaluation for Large Language Models
Shiguo Lian, Kaikai Zhao, Xuejiao Lei, Bikun Yang, Wenjing Zhang 0006, Kai Wang 0012, Zhaoxiang Liu
NLPCC (3)7
2024 Optimized Conversational Gesture Generation with Enhanced Motion Feature Extraction and Cascaded Generator
Xiang Wang 0018, Yifeng Peng, Zhaoxiang Liu, Shijie Dong, Ruitao Liu, Kai Wang 0012, Shiguo Lian
NLPCC (3)6
2024 Hybrid attention transformer with re-parameterized large kernel convolution for image super-resolution
Zhicheng Ma, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian
Image Vis. Comput.3
2020 A survey on face data augmentation for the training of deep neural networks
Xiang Wang 0018, Kai Wang 0012, Shiguo Lian
Neural Comput. Appl.2
2019 Real-Time 3D Object Detection and Tracking in Monocular Images of Cluttered Environment
Guoguang Du 0001, Kai Wang 0012, Yibing Nan, Shiguo Lian
ICIG (2)2
2019 A Unified Framework for Mutual Improvement of SLAM and Semantic Segmentation
abstract
This paper presents a novel framework for simultaneously implementing localization and segmentation, which are two of the most important vision-based tasks for robotics. While the goals and techniques used for them were considered to be different previously, we show that by making use of the intermediate results of the two modules, their performance can be enhanced at the same time. Our framework is able to handle both the instantaneous motion and long-term changes of instances in localization with the help of the segmentation result, which also benefits from the refined 3D pose information. We conduct experiments on various datasets, and prove that our framework works effectively on improving the precision and robustness of the two tasks and outperforms existing localization and segmentation algorithms.
Kai Wang 0012, Yimin Lin, Luowei Wang, Liming Han, Minjie Hua, Xiang Wang 0018, Shiguo Lian, Bill Huang
ICRA1
2019 Towards More Realistic Human-Robot Conversation: A Seq2Seq-based Body Gesture Interaction System
abstract
This paper presents a novel system that enables intelligent robots to exhibit realistic body gestures while communicating with humans. The proposed system consists of a listening model and a speaking model used in corresponding conversational phases. Both models are adapted from the sequence-to-sequence (seq2seq) architecture to synthesize body gestures represented by the movements of twelve upper-body keypoints. All the extracted 2D keypoints are firstly 3D-transformed, then rotated and normalized to discard irrelevant information. Substantial videos of human conversations from Youtube are collected and preprocessed to train the listening and speaking models separately, after which the two models are evaluated using metrics of mean squared error (MSE) and cosine similarity on the test dataset. The tuned system is implemented to drive a virtual avatar as well as Pepper, a physical humanoid robot, to demonstrate the improvement on conversational interaction abilities of our method in practice.
Minjie Hua, Fuyuan Shi, Yibing Nan, Kai Wang 0012, Shiguo Lian
IROS4
2019 Progressive sketching with instant previewing
Kai Wang 0012, Jianmin Zheng, Seah Hock Soon
Comput. Graph.1
2018 Enhancing Sketching and Sculpting for Shape Modeling
abstract
Sketch-based modeling uses freeform strokes as basic modeling metaphor and provides an intuitive way for shape modeling, for instance, for cyberworlds. This paper presents a new method to enhance sketch-based modeling. The core idea of the method is to enhance the sketching process by allowing the user to iteratively sketch to progressively create initial shapes that interpolate the sketched strokes. This process considers all the sketches and the up-to-date constructed 3D shape, which enables the user to be aware of the shape of the sketched model. The key underlying technique that supports this process is a novel surface construction algorithm, which generates 3D triangular mesh models with gradual shape changes during iterative sketching. Experiments demonstrate that the presented method can allow users to intuitively and flexibly create and edit 3D models even with complex topology, which is usually difficult in existing sketch-based modeling systems.
Kai Wang 0012, Jianmin Zheng, Seah Hock Soon
CW1
2010 Reference Plane Assisted Sketching Interface for 3D Freeform Shape Design
abstract
This paper presents a sketch-based modeling system with auxiliary planes as references for 3D freeform shape design. The user first creates a rough 3D model of arbitrary topology by sketching some contours of the model. Then the user can use sketching to perform deformation, extrusion, etc, to edit the model. To regularize and interpret the user's inputs properly, we introduce some rules for the strokes into the system, which are based on both the semantic meaning of the sketched strokes and human psychology. Unlike other sketching systems, all the creation and editing operations in the presented system are performed with reference to some auxiliary planes that are automatically constructed based on the user’s sketches or default settings. The use of reference planes provides a heuristic solution to the problem of ambiguity of 2D interface for modeling in 3D space. Examples demonstrate that the presented system can allow the user to intuitively and intelligently create and edit 3D models even with complex topology, which is usually difficult in other similar sketch-based modeling systems.
Kai Wang 0012, Jianmin Zheng, Seah Hock Soon
CW1