EDBT 2026 Demo / reviewers in the wild / expert
Yong-Jin Liu 0001
dblp:03/6872-1 · also Yongjin Liu 0001
· DBLP profile ↗
216ranked-venue papers
43as first author
118since 2021 · last 2026
0000-0001-5774-1916ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 137 · 28 first-author · 67 since 2021Artificial intelligence and machine learning · 83 · 8 first-author · 60 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 6 first-author · 5 since 2021Systems, architecture and hardware · 12 · 8 since 2021Human-computer interaction and ubiquitous computing · 8 · 6 since 2021Theory of computation · 6 · 5 first-authorComputer networks · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cybersickness Exploration for Different VR Tasks Under Variable Rendering ConditionsabstractCybersickness is a major challenge in virtual reality (VR), adversely affecting user comfort and usability. While prior research has examined visual and system factors, the impact of different VR tasks under varied rendering conditions remains underexplored. To address this, we investigated three interaction tasks—navigation, selection, and manipulation—together with two rendering parameters: center area radius (CAR) and peripheral resolution (PR). Cybersickness severity and frequency were assessed using Simulator Sickness Questionnaire (SSQ) scores and electroencephalography (EEG) data. Results show that task type significantly affects cybersickness severity, with CAR playing a critical role. Moreover, α-band power spectral density (PSD) strongly correlates with SSQ scores, suggesting its potential as a biomarker for cybersickness. Neural responses also exhibited temporal delays compared to subjective reports, offering insights into the mechanisms of cybersickness. These findings advance understanding of task- and rendering-related influences, informing more effective prediction and mitigation strategies in VR system design. Peike Wang, Jian Wu 0033, Lili Wang 0006, Yong-Jin Liu 0001 |
Int. J. Hum. Comput. Interact. | 5 |
| 2026 | Positive or Negative? Exploring the Effect of Self-Motion Speed on Flow Experience Based on an Educational GameabstractActivities conducted in games utilizing immersive virtual reality (VR) environment commonly require users to actively navigate from a first-person perspective. The sense of self-motion from active navigation may contribute to flow experience, but it also increases the risk of VR sickness. Clarifying their underlying relationships can help optimize the design of VR games. To explore this issue, this paper developed an experimental VR environment supporting active navigation at different speeds and subsequently performed two experiments. Experiment 1 first examined the effect of self-motion speed on flow experience and VR sickness. Experiment 2 examined the potential moderating effect of VR sickness susceptibility in the relationship between self-motion speed and flow experience. The results showed that the effect of self-motion speed on flow experience was moderated by VR sickness susceptibility. For participants with low susceptibility, self-motion speed had a significantly positive effect on flow experience, and for participants with high susceptibility, the effect was not significant. Chao Zhou 0012, Yong-Jin Liu 0001, Yulong Bian |
Int. J. Hum. Comput. Interact. | 2 |
| 2026 | Dyadic Imitation Modeling With Lag-Aware Dual-Stream Framework for Autism ClassificationabstractThis paper introduces a novel lag-aware dual-stream (LADS) framework and a carefully curated dual-view video dataset for automatic Autism Spectrum Disorder (ASD) classification through imitation tasks. In contrast to prior single-view approaches that overlook the interactive dynamics of imitation, our dataset is the first to capture synchronized experimenter-child interactions with rich pose and motion features. Building on this data, the LADS framework explicitly learns the temporal alignment between the experimenter’s demonstration and the child’s imitative response. A Lag-Aware Alignment module uses constrained cross-attention to compute an adaptive time warping and extract per-frame lag feature, revealing delays in the child’s imitation. Additionally, a lightweight diffusion-based regularizer enforces representation consistency by denoising perturbed child features conditioned on the experimenter’s motion, improving generalization. We then achieve the classification of ASD versus Typical Development (TD) behavior by integrating the aligned dual-stream features, imitation lag, and action discrepancy within an attention-pooling classifier. Experiments on our dual-view imitation dataset show that LADS significantly outperforms conventional single-stream models and a recent dyadic transformer baseline, achieving state-of-the-art classification accuracy. The results demonstrate the importance of modeling interpersonal timing in social behavior analysis. Our work provides a new, public dataset and a computational tool for interdisciplinary research, bridging computer vision and psychological studies of autism. Both dataset and code will be made publicly available. Wenqi Ji, Ying Guo 0004, Kaiyun Li, Yong-Jin Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | Internal State Estimation in Crowds via Active Information GatheringabstractAccurately estimating human internal states, such as personality traits or behavioral patterns, is critical for enhancing the effectiveness of human–robot interaction, particularly in multi-agent settings. These insights are key in applications ranging from social navigation to autism diagnosis. However, prior methods are limited by scalability and passive observation, making real-time estimation in complex, multi-human settings difficult. In this work, we propose a practical method for active human personality estimation in crowds, with a focus on applications related to Autism Spectrum Disorder (ASD). Our method combines a personality-conditioned behavior model, based on the Eysenck 3-Factor theory, with an active robot information-gathering policy that triggers human behaviors through a receding-horizon planner. The robot’s belief about human personality is then updated via Bayesian inference. We demonstrate the effectiveness of our approach through proof-of-concept studies in simulation, user studies with typical adults, and preliminary experiments involving participants with ASD. Our results show that our method can scale to tens of humans and reduce personality estimation error by 29.2% and uncertainty by 79.9% in simulation compared to the passive baseline. User studies with typical adults confirm the method’s ability to generalize across complex personality distributions. Additionally, we explore its application in autism-related scenarios, demonstrating that the method can identify the difference between neurotypical and autistic behavior. The results suggest that our framework could serve as a foundation for future ASD-specific applications. Xuebo Ji, Zherong Pan, Xifeng Gao, Lei Yang 0048, Xinxin Du, Kaiyun Li, Yong-Jin Liu 0001, Wenping Wang 0001, Changhe Tu, Jia Pan 0001 |
ACM Trans. Hum. Robot Interact. | 7 |
| 2026 | Progressive Orthodontic Motion Planning Based on Hierarchical Diffusion TransformerabstractOrthodontic motion planning plays a crucial role in digital orthodontics by predicting tooth motion sequences to assist dentists in formulating treatment plans efficiently. Most prior work generates the entire intermediate tooth motion sequence given the initial and target tooth alignments. In practice, only the initial alignment of the patient is obtained. However, no existing method can predict the complete motion sequence using only the initial tooth alignment. To address this gap, we propose OrthoDiff, a novel target-free framework that uses only initial tooth alignment through a progressive generation strategy. This strategy generates tooth motion sequences by decomposing the entire motion sequence into multi-level motions, progressively constraining the inference space and reducing the complexity of target-free planning from coarse to fine. Moreover, we design a hierarchical diffusion transformer as the backbone of OrthoDiff, which treats tooth alignment as a sequence of tooth tokens and fully leverages the topological prior knowledge of the dental model. Through extensive evaluations, we demonstrate that our method significantly outperforms state-of-the-art techniques in target-free tooth motion generation. Ablation studies further confirm the efficacy of key components in our network design. Meanwhile, we also achieve state-of-the-art results in tooth target alignment prediction, benefiting from our framework. The code and data will be publicly available at https://github.com/Intelligent-Orthodontics/OrthoDiff.github.io. Yeying Fan, Yuanfeng Zhou, Guangshun Wei, Zhiming Cui 0001, Yiran Shen 0001, Yong-Jin Liu 0001, Wenping Wang 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2026 | A Text-to-3D Framework for Joint Generation of CG-Ready Humans and Compatible GarmentsabstractCreating detailed 3D human avatars with fitted garments traditionally requires specialized expertise and labor-intensive workflows. While recent advances in generative AI have enabled text-to-3D human and clothing synthesis, existing methods fall short in offering accessible, integrated pipelines for generating CG-ready 3D avatars with physically compatible outfits; here we use the term CG-ready for models following a technical aesthetic common in computer graphics (CG) and adopt standard CG polygonal meshes and strands representations (rather than radiance field representations like NeRF and 3DGS) that can be directly integrated into conventional CG pipelines and support downstream tasks such as physical simulation. To bridge this gap, we introduce Tailor, an integrated text-to-3D framework that generates high-fidelity, customizable 3D avatars dressed in simulation-ready garments. Tailor consists of three stages. (1) Semantic Parsing: we employ a large language model to interpret textual descriptions and translate them into parameterized human avatars and semantically matched garment templates. (2) Geometry-Aware Garment Generation: we propose topology-preserving deformation with novel geometric losses to generate body-aligned garments under text control. (3) Consistent Texture Synthesis: we propose a novel multi-view diffusion process optimized for garment texturing, which enforces view consistency, preserves photorealistic details, and optionally supports symmetric texture generation common in garments. Through comprehensive quantitative and qualitative evaluations, we demonstrate that Tailor outperforms state-of-the-art methods in fidelity, usability, and diversity. Zhiyao Sun, Yu-Hui Wen, Ho-Jui Fang, Matthieu Lin, Tian Lv, Yong-Jin Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | Weighted Poisson-disk Resampling on Large-Scale Point CloudsabstractFor large-scale point cloud processing, resampling takes the important role of controlling the point number and density while keeping the geometric consistency. However, current methods cannot balance such different requirements. Particularly with large-scale point clouds, classical methods often struggle with decreased efficiency and accuracy. To address such issues, we propose a weighted Poisson-disk (WPD) resampling method to improve the usability and efficiency for the processing. We first design an initial Poisson resampling with a voxel-based estimation strategy. It is able to estimate a more accurate radius of the Poisson-disk while maintaining high efficiency. Then, we design a weighted tangent smoothing step to further optimize the Voronoi diagram for each point. At the same time, sharp features are detected and kept in the optimized results with isotropic property. Finally, we achieve a resampling copy from the original point cloud with the specified point number, uniform density, and high-quality geometric consistency. Experiments show that our method significantly improves the performance of large-scale point cloud resampling for different applications, and provides a highly practical solution. Xianhe Jiao, Chenlei Lv, Junli Zhao, Ran Yi 0002, Yu-Hui Wen, Zhenkuan Pan 0001, Zhongke Wu, Yong-Jin Liu 0001 |
AAAI | 8 |
| 2025 | DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing ConstraintsabstractRecent advances in large language model assistants have made them indispensable, raising significant concerns over managing their safety. Automated red teaming offers a promising alternative to the labor-intensive and error-prone manual probing for vulnerabilities, providing more consistent and scalable safety evaluations. However, existing approaches often compromise diversity by focusing on maximizing attack success rate. Additionally, methods that decrease the cosine similarity from historical embeddings with semantic diversity rewards lead to novelty stagnation as history grows. To address these issues, we introduce DiveR-CT, which relaxes conventional constraints on the objective and semantic reward, granting greater freedom for the policy to enhance diversity. Our experiments demonstrate DiveR-CT's marked superiority over baselines by 1) generating data that perform better in various diversity metrics across different attack success rate levels, 2) better-enhancing resiliency in blue team models through safety tuning based on collected data, 3) allowing dynamic control of objective weights for reliable and controllable attack success rates, and 4) reducing susceptibility to reward overoptimization. Overall, our method provides an effective and efficient approach to LLM red teaming, accelerating real-world deployment. WARNING: This paper contains examples of potentially harmful text. Andrew Zhao, Quentin Xu, Matthieu Lin, Shenzhi Wang, Yong-Jin Liu 0001, Zilong Zheng, Gao Huang 0001 |
AAAI | 5 |
| 2025 | VAction: A Lightweight and Integrated VR Training System for Authentic Film-Shooting Experience
Che Qu, Minjing Yu, Chao Zhou 0012, Yuntao Wang 0001, Yu-Hui Wen, Yuanchun Shi, Yong-Jin Liu 0001 |
CHI | 8 |
| 2025 | Sketch-Guided Scene-Level Image Editing with Diffusion Models
Ran Zuo, Haoxiang Hu, Xiaoming Deng 0001, Yaokun Li, Yukun Lai, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
CVM (2) | 7 |
| 2025 | StdGEN: Semantic-Decomposed 3D Character Generation from Single ImagesabstractWe present StdGEN, an innovative pipeline for generating semantically decomposed high-quality 3D characters from single images, enabling broad applications in virtual reality, gaming, and filmmaking, etc. Unlike previous methods which struggle with limited decomposability, unsatisfactory quality, and long optimization times, StdGEN features decomposability, effectiveness and efficiency; i.e., it generates intricately detailed 3D characters with separated semantic components such as the body, clothes, and hair, in three minutes. At the core of StdGEN is our proposed Semantic-aware Large Reconstruction Model (S-LRM), a transformer-based generalizable model that jointly reconstructs geometry, color and semantics from multi-view images in a feed-forward manner. A differentiable multilayer semantic surface extraction scheme is introduced to acquire meshes from hybrid implicit fields reconstructed by our S-LRM. Additionally, a specialized efficient multi-view diffusion model and an iterative multi-layer surface refinement module are integrated into the pipeline to facilitate high-quality, decomposable 3D character generation. Extensive experiments demonstrate our state-of-theart performance in 3D anime character generation, surpassing existing baselines by a significant margin in geometry, texture and decomposability. StdGEN offers ready-touse semantic-decomposed 3D characters and enables flexible customization for a wide range of applications. Project page: https://stdgen.github.io Yanning Zhou 0003, Wang Zhao 0001, Zhongkai Wu, Kaiwen Xiao, Wei Yang 0013, Yong-Jin Liu 0001, Xiao Han 0011 |
CVPR | 7 |
| 2025 | Rectified Diffusion Guidance for Conditional GenerationabstractClassifier-Free Guidance (CFG), which combines the conditional and unconditional score functions with two coefficients summing to one, serves as a practical technique for diffusion model sampling. Theoretically, however, denoising with CFG cannot be expressed as a reciprocal diffusion process, which may consequently leave some hidden risks during use. In this work, we revisit the theory behind CFG and rigorously confirm that the improper configuration of the combination coefficients (i.e., the widely used summing-to-one version) brings about expectation shift of the generative distribution. To rectify this issue, we propose ReCFG1with a relaxation on the guidance coefficients such that denoising with ReCFG strictly aligns with the diffusion theory. We further show that our approach enjoys a closed-form solution given the guidance strength. That way, the rectified coefficients can be readily pre-computed via traversing the observed data, leaving the sampling speed barely affected. Empirical evidence on real-world data demonstrate the compatibility of our post-hoc design with existing state-of-the-art diffusion models, including both class-conditioned ones (e.g., EDM2 on ImageNet) and text-conditioned ones (e.g., SD3 on CC12M), without any retraining. Code is available at https://github.com/thuxmf/recfg. Mengfei Xia, Nan Xue 0001, Yujun Shen, Ran Yi 0002, Tieliang Gong, Yong-Jin Liu 0001 |
CVPR | 6 |
| 2025 | TeethGenerator: A Two-Stage Framework for Paired Pre- and Post-Orthodontic 3D Dental Data GenerationabstractDigital orthodontics represents a prominent and critical application of computer vision technology in the medical field. So far, the labor-intensive process of collecting clinical data, particularly in acquiring paired 3D orthodontic teeth models, constitutes a crucial bottleneck for developing tooth arrangement neural networks. Although numerous general 3D shape generation methods have been proposed, most of them focus on single-object generation and are insufficient for generating anatomically structured teeth models, each comprising 24-32 segmented teeth. In this paper, we propose TeethGenerator, a novel two-stage framework designed to synthesize paired 3D teeth models pre- and post-orthodontic, aiming to facilitate the training of downstream tooth arrangement networks. Specifically, our approach consists of two key modules: (1) a teeth shape generation module that leverages a diffusion model to learn the distribution of morphological characteristics of teeth, enabling the generation of diverse post-orthodontic teeth models; and (2) a teeth style generation module that synthesizes corresponding pre-orthodontic teeth models by incorporating desired styles as conditional inputs. Extensive qualitative and quantitative experiments demonstrate that our synthetic dataset aligns closely with the distribution of real orthodontic data, and promotes tooth alignment performance significantly when combined with real data for training. The code and dataset are available at https://github.com/lcshhh/teeth_generator. Changsong Lei, Yaqian Liang, Shaofeng Wang, Jiajia Dai, Yong-Jin Liu 0001 |
ICCV | 5 |
| 2025 | GauSurfaceAvatar: A Realistic Human Head Model with Variable Texture Based on 2D Gaussiansabstract3D facial reconstruction plays a crucial role in virtual reality and entertainment. Impressive rendering and animation effects have been achieved through recent advances. Existing Gaussian avatar models generate diverse expressions, but the corresponding texture changes are not so satisfactory, and their geometric structure often falls short. In response to these challenges, we propose a 3D head avatar with remarkable geometry, which can control expression variations and enable facial texture information to change along with expressions. To achieve this effect, we combine the 2D Gaussian field with the facial parametric model, use the mesh to drive Gaussian field, design a fine-tuning field for the mouth area to fit distorted expressions, and simultaneously design a color variation module to simulate changes in facial wrinkles and skin. Experiments demonstrate that our avatar model exhibits excellent performance in both appearance details and geometric shapes. Lijie Geng, Junli Zhao, Lin Gao 0004, Ran Yi 0002, Fuqing Duan, Zhenkuan Pan 0001, Yong-Jin Liu 0001 |
ICME | 7 |
| 2025 | CHARM: Control-point-based 3D Anime Hairstyle Auto-Regressive ModelingabstractWe present CHARM, a novel parametric representation and generative framework for anime hairstyle modeling. While traditional hair modeling methods focus on realistic hair using strand-based or volumetric representations, anime hairstyle exhibits highly stylized, piecewise-structured geometry that challenges existing techniques. Existing works often rely on dense mesh modeling or hand-crafted spline curves, making them inefficient for editing and unsuitable for scalable learning. CHARM introduces a compact, invertible control-point-based parameterization, where a sequence of control points represents each hair card, and each point is encoded with only five geometric parameters. This efficient and accurate representation supports both artist-friendly design and learning-based generation. Built upon this representation, CHARM introduces an autoregressive generative framework that effectively generates anime hairstyles from input images or point clouds. By interpreting anime hairstyles as a sequential “hair language”, our autoregressive transformer captures both local geometry and global hairstyle topology, resulting in high-fidelity anime hairstyle creation. To facilitate both training and evaluation of anime hairstyle generation, we construct AnimeHair, a large-scale dataset of 37K high-quality anime hairstyles with separated hair cards and processed mesh data. Extensive experiments demonstrate state-of-the-art performance of CHARM in both reconstruction accuracy and generation quality, offering an expressive and scalable solution for anime hairstyle modeling. Project page: https://hyzcluster.github.io/charm Yanning Zhou 0003, Wang Zhao 0001, Jingwen Ye, Yushi Bai, Kaiwen Xiao, Yong-Jin Liu 0001, Zhongqian Sun, Wei Yang 0032 |
SIGGRAPH Asia | 7 |
| 2025 | SketchGPT: A Sketch-based Multimodal Interface for Application-Agnostic LLM Interaction
Cangjun Gao, Yaxian Shan, Haoxiang Hu, Qingkun Li, Xiaoming Deng 0001, CuiXia Ma, Yukun Lai, Yong-Jin Liu 0001, Feng Tian 0001, Guozhong Dai, Hongan Wang |
UIST | 9 |
| 2025 | Sketch123: Multi-spectral channel cross attention for sketch-based 3D generation via diffusion models
Zhentong Xu, Long Zeng 0001, Junli Zhao, Baodong Wang, Zhenkuan Pan 0001, Yong-Jin Liu 0001 |
Comput. Aided Des. | 6 |
| 2025 | A multi-view projection-based object-aware graph network for dense captioning of point clouds
Zijing Ma, Aihua Mao, Shuyi Wen, Ran Yi 0002, Yong-Jin Liu 0001 |
Comput. Graph. | 6 |
| 2025 | Generalized Robot Vision-Language Model via Linguistic Foreground-Aware Contrast
Kangcheng Liu, Xiaodong Han, Yong-Jin Liu 0001, Baoquan Chen |
Int. J. Comput. Vis. | 4 |
| 2025 | Correction: Generalized Robot Vision-Language Model via Linguistic Foreground-Aware Contrast
Kangcheng Liu, Xiaodong Han, Yong-Jin Liu 0001, Baoquan Chen |
Int. J. Comput. Vis. | 4 |
| 2025 | Self-Referencing Agents for Unsupervised Reinforcement Learning
Andrew Zhao, Erle Zhu, Rui Lu 0001, Matthieu Lin, Yong-Jin Liu 0001, Gao Huang 0001 |
Neural Networks | 5 |
| 2025 | Corrigendum to "Self-Referencing agents for unsupervised reinforcement learning" [Neural Networks Volume 188, August 2025, 107448]
Andrew Zhao, Erle Zhu, Rui Lu 0001, Matthieu Lin, Yong-Jin Liu 0001, Gao Huang 0001 |
Neural Networks | 5 |
| 2025 | General 3D Vision-Language Model With Fast Rendering and Pre-Training Vision-Language AlignmentabstractCurrent prevailing vision-language models have achieved remarkable progress in 3D scene understanding while trained in the closed-set setting and with full labels. The major bottleneck for the current robot 3D scene recognition approach for robotic applications is that these models do not have the capacity to recognize any unseen novel classes beyond the training categories in diverse real-world robot applications such as robot manipulation as well as robot navigation. In the meantime, current state-of-the-art 3D scene understanding approaches primarily require a large number of high-quality labels to train neural networks, which merely perform well in a fully supervised manner. Therefore, we are in urgent need of a framework that can simultaneously be applicable to both 3D point cloud segmentation and detection, particularly in the circumstances where the labels are rather scarce. This work presents a generalized and straightforward framework for dealing with 3D scene understanding when the labeled scenes are quite limited. To extract knowledge for novel categories from the pre-trained vision-language models, we propose a hierarchical feature-aligned pre-training and knowledge distillation strategy to extract and distill meaningful information from large-scale vision-language models, which helps benefit the open-vocabulary scene understanding tasks. To leverage the boundary information, we propose a novel energy-based loss with boundary awareness benefiting from the region-level boundary predictions. To encourage latent instance discrimination and to guarantee efficiency, we propose the unsupervised region-level semantic contrastive learning scheme for point clouds, using confident predictions of the neural network to discriminate the intermediate feature embeddings at multiple stages. In the limited reconstruction case, our proposed approach, termed WS3D++, ranks 1st on the large-scale ScanNet benchmark on both the task of semantic segmentation and instance segmentation. Also, our proposed WS3D++ achieves state-of-the-art data-efficient learning performance on the other large-scale real-scene indoor and outdoor datasets S3DIS and SemanticKITTI. Extensive experiments with both indoor and outdoor scenes demonstrated the effectiveness of our approach in both data-efficient learning and open-world few-shot learning. Kangcheng Liu, Yong-Jin Liu 0001, Baoquan Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | LibEER: A Comprehensive Benchmark and Algorithm Library for EEG-Based Emotion RecognitionabstractEEG-based emotion recognition (EER) has gained significant attention due to its potential for understanding and analyzing human emotions. While recent advancements in deep learning techniques have substantially improved EER, the field lacks a convincing benchmark and comprehensive open-source libraries. This absence complicates fair comparisons between models and creates reproducibility challenges for practitioners, which collectively hinder progress. To address these issues, we introduce LibEER, a comprehensive benchmark and algorithm library designed to facilitate fair comparisons in EER. LibEER carefully selects popular and powerful baselines, harmonizes key implementation details across methods, and provides a standardized codebase in PyTorch. By offering a consistent evaluation framework with standardized experimental settings, LibEER enables unbiased assessments of seventeen representative deep learning models for EER across the six most widely used datasets. Additionally, we conduct a thorough, reproducible comparison of model performance and efficiency, providing valuable insights to guide researchers in the selection and design of EER models. Moreover, we make observations and in-depth analysis on the experiment results and identify current challenges in this community. We hope that our work will not only lower entry barriers for newcomers to EEG-based emotion recognition but also contribute to the standardization of research in this domain, fostering steady development. The library and source code are publicly available athttps://github.com/XJTU-EEG/LibEER. Huan Liu 0012, Shusen Yang, Yuzhe Zhang 0003, Mengze Wang, Fanyu Gong, Chengxi Xie, Guanjian Liu, Zejun Liu, Yong-Jin Liu 0001, Bao-Liang Lu, Dalin Zhang 0001 |
IEEE Trans. Affect. Comput. | 9 |
| 2025 | Temporal-Noise-Aware Neural Networks for Suicidal Ideation Prediction Using Physiological DataabstractThe robust generalization of deep learning models in the presence of inherent noise remains a significant challenge, especially when labels are ambiguous due to their subjective nature and noise is indiscernible in natural settings. In this article, we address a specific and important scenario of monitoring suicidal ideation (SI), where time-series data, such as galvanic skin response (GSR) and photoplethysmography (PPG), are susceptible to such noise. Current methods predominantly focus on image and text data or address artificially introduced noise, neglecting the complexities of natural noise in time-series analysis. To tackle this, we introduce a novel neural network model tailored for analyzing noisy physiological time-series data, named DBN_ConvNet, which integrates advanced encoding techniques with confidence learning training to enhance prediction performance. Another main contribution of our work is the collection of a specialized dataset of GSR and PPG signals derived from real-world environments for SI prediction. By employing this dataset, our DBN_ConvNet achieves a prediction accuracy of 76.67% and an F1 score of 0.74 in a binary classification task, outperforming state-of-the-art methods. Furthermore, comprehensive evaluations have been conducted on three other well-known public datasets with artificially introduced noise to test the DBN_ConvNet’s capabilities rigorously. These tests consistently demonstrated DBN_ConvNet’s superior performance by achieving an improvement of more than 10% in both accuracy and F1 score compared to the baseline methods. Niqi Liu, Fang Liu 0035, Xinxin Du, Yezhi Shu, Xu Liu 0006, Guozhen Zhao, Wenting Mu, Yong-Jin Liu 0001 |
IEEE Trans. Comput. Soc. Syst. | 9 |
| 2025 | Exploring the Influence of Profile Picture Styles on Empathy and Identity Recognition in Social MediaabstractEmpathy and identity recognition are two core social interaction factors that greatly affect efficiency and effectiveness. The selection of a profile picture in social media is crucial as it serves as a visual representation of virtual identity. It not only reflects the user's personality but also influences the level of empathy and identification that others feel toward them. However, the potential impact of profile pictures with different styles (e.g., real faces, cartoon faces, and landscape images) on empathy and identity recognition is still unclear. To explore its effects, a controlled laboratory experiment and an ecological online experiment were conducted. Participants were shown a picture each time, informed to imagine interacting with the person using it as his/her profile picture, and instructed to rate an item from the basic empathy scale (BES) based on it. After rating all pictures, users then completed an identity recognition task. Results show that participants’ empathy scores for users with cartoon or real face profile pictures are greater than those with landscape profile pictures. In addition, participants performed better in identity recognition for users with real face or landscape images as profile pictures than for those with cartoon face profile pictures. Moreover, users of social media often make social categorizations (i.e., in-group/out-group categorization) based on the social identities expressed by their profile pictures. Our results also indicate that the affective empathy scale rating score was positively associated with the degree to which users of the corresponding profile pictures were categorized as in-group members. Minjing Yu, Xinge Liu, Chao Zhou 0012, Xinxin Du, Jenny Sheng, Yong-Jin Liu 0001 |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2025 | DRC: Discrete Representation Classifier With Salient Features via Fixed-PrototypeabstractImage classification models including convolutional neural networks (CNN) and vision transformers (ViT) commonly employ a fully connected (FC) layer as the classifier. However, the fully connected nature of FC brings large amounts of weight parameters, limits the efficiency of inference, tends to over-fit the training data, and struggles to learn distinct class weights. To solve these problems, we propose a discrete representation classifier (DRC), a generic parameter-free classifier that offers efficiency, robustness, and more discriminative categorization. Specifically, the DRC discards numerous unimportant features and focuses solely on the salient features which are reinforced during training and presented in short discrete form during inference. Unlike the way of learning pseudo-prototypes (weights) from data laden with complex patterns and noises in FC, the DRC introducing discriminative fixed-prototypes which are almost uniformly distributed across the high-dimensional feature space, thus helps the model to learn more distinct boundaries between categories. Further leveraging the advantage of DRC’s focus on salient features, we propose Salient-CAM, which is able to locate the most important region in image without the need for weighting feature maps. The experiments demonstrate that simply replacing the model’s classifier from FC to DRC can lead to a significant acceleration in the whole model’s inference and a more robust classification. Additionally, the proposed Salient-CAM exhibits excellent object localization ability in complex natural scenes. Qinglei Li, Qi Wang 0079, Yongbin Qin, Xingcai Wu, Shiming Chen 0002, Wu Liu 0005, Yong-Jin Liu 0001, Jiebo Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | Robust 3D Visual Question Answering via Bias LearningabstractVisual question answering (VQA) tasks have witnessed significant advancements in recent years. So far, enhancing the robustness of models on diverse datasets and improving their performance in3D environmentsremains a challenging research direction. In this paper, we propose a high-performance framework called Bias3D-VQA for 3D-VQA based on generative adversarial networks via bias learning, addressing the inherent biases that arise from the model’s dependency on dataset-specific patterns or tendencies during training. Such biases often lead the model to focus on more frequently occurring but incorrect answers. Our framework comprises a target model, a bias model, and a generative adversarial component. In each training iteration, we employ an alternating training approach for the target and bias models. When training the bias model, fake point cloud data is generated from random noise, and then we accumulate biases present in language modality and various modules through adversarial training. When training the target model, both the question and the 3D point cloud are inputted into the bias model simultaneously, and the output of the bias model is utilized to correct the loss of the target model. Our approach(Bias3D-VQA) is the first to focus on enhancing model robustness by addressing diverse biases in the 3D-VQA domain. Our target model demonstrates superior performance compared to state-of-the-art models, showing significant improvements in classification accuracy and text generation quality. Notably, in the metrics such as EM@1 and CIDEr, our model even surpasses some pre-trained models with large additional datasets. The source code is available athttps://github.com/coderr727/bias_3DQA Aihua Mao, Shuyi Wen, Ran Yi 0002, Yong-Jin Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Autonomous Tomato Harvesting With Top-Down Fusion Network for Limited DataabstractUsing robots for tomato truss harvesting represents a promising approach to agricultural production. However, incomplete acquisition of perception information and clumsy operations often result in low harvest success rates or crop damage. To address this issue, we designed a new method for tomato truss perception, an autonomous harvesting method, and a novel circular rotary cutting end-effector. The robot performs object detection and keypoint detection on tomato trusses using the proposed Top-down Fusion Network, making decisions on suitable targets for harvesting based on phenotyping and pose estimation. The designed end-effector moves gradually from the bottom up to wrap around the tomato truss, cutting the peduncle to complete the harvest. Experiments conducted in real-world scenarios for robotic perception and autonomous harvesting of tomato trusses show that the proposed method increases accuracy by up to 11.42% and 22.29% for complete and limited dataset conditions, compared to baseline models. Furthermore, we have implemented an automatic tomato harvesting system based on TDFNet, which reaches an average harvest success rate of 89.58% in the greenhouse. Xingxu Li, Yiheng Han, Nan Ma 0012, Yong-Jin Liu 0001, Jia Pan 0001, Siyi Zheng |
IEEE Trans. Robotics | 4 |
| 2025 | PCKRF: Point Cloud Completion and Keypoint Refinement With Fusion Data for 6D Pose EstimationabstractSome robust point cloud registration approaches with controllable pose refinement magnitude, such as ICP and its variants, are commonly used to improve 6D pose estimation accuracy. However, the effectiveness of these methods gradually diminishes with the advancement of deep learning techniques and the enhancement of initial pose accuracy, primarily due to their lack of specific design for pose refinement. In this paper, we propose Point Cloud Completion and Keypoint Refinement with Fusion Data (PCKRF), a new pose refinement pipeline for 6D pose estimation. The pipeline consists of two steps. First, it completes the input point clouds via a novel pose-sensitive point completion network. The network uses both local and global features with pose information during point completion. Then, it registers the completed object point cloud with the corresponding target point cloud by our proposed Color supported Iterative KeyPoint (CIKP) method. The CIKP method introduces color information into registration and registers a point cloud around each keypoint to increase stability. The PCKRF pipeline can be integrated with existing popular 6D pose estimation methods, such as the full flow bidirectional fusion network, to further improve their pose estimation accuracy. Experiments demonstrate that our method exhibits superior stability compared to existing approaches when optimizing initial poses with relatively high precision. Notably, the results indicate that our method effectively complements most existing pose estimation techniques, leading to improved performance in most cases. Furthermore, our method achieves promising results even in challenging scenarios involving textureless and symmetrical objects. Yiheng Han, Irvin Haozhe Zhan, Long Zeng 0001, Yu-Ping Wang 0001, Ran Yi 0002, Minjing Yu, Matthieu Lin, Jenny Sheng, Yong-Jin Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2025 | Multimodal Contrastive Learning for Cybersickness Recognition Using Brain Connectivity Graph RepresentationabstractCybersickness significantly impairs user comfort and immersion in virtual reality (VR). Effective identification of cybersickness leveraging physiological, visual, and motion data is a critical prerequisite for its mitigation. However, current methods primarily employ direct feature fusion across modalities, which often leads to limited accuracy due to inadequate modeling of inter-modal relationships. In this paper, we propose a multimodal contrastive learning method for cybersickness recognition. First, we introduce Brain Connectivity Graph Representation (BCGR), an innovative graph-based representation that captures cybersickness-related connectivity patterns across modalities. We further develop three BCGR instances: E-BCGR, constructed based on EEG signals; MV-BCGR, constructed based on video and motion data; and S-BCGR, obtained through our proposed standardized decomposition algorithm. Then, we propose a connectivity-constrained contrastive fusion module, which aligns E-BCGR and MV-BCGR into a shared latent space via graph contrastive learning while utilizing S-BCGR as a connectivity constraint to enhance representation quality. Moreover, we construct a multimodal cybersickness dataset comprising synchronized EEG, video, and motion data collected in VR environments to promote further research in this domain. Experimental results demonstrate that our method outperforms existing state-of-the-art methods across four critical evaluation metrics: accuracy, sensitivity, specificity, and the area under the curve. Source code: https://github.com/PEKEW/cybersickness-bcgr. Peike Wang, Ziteng Wang 0002, Yong-Jin Liu 0001, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | Indoor Scene Reconstruction With Fine-Grained Details Using Hybrid Representation and Normal Prior EnhancementabstractThe reconstruction of indoor scenes from multi-view RGB images is challenging due to the coexistence of flat and texture-less regions alongside delicate and fine-grained regions. Recent methods leverage neural radiance fields aided by predicted surface normal priors to recover the scene geometry. These methods excel in producing complete and smooth results for floor and wall areas. However, they struggle to capture complex surfaces with high-frequency structures due to the inadequate neural representation and the inaccurately predicted normal priors. This work aims to reconstruct high-fidelity surfaces with fine-grained details by addressing the above limitations. To improve the capacity of the implicit representation, we propose a hybrid architecture to represent low-frequency and high-frequency regions separately. To enhance the normal priors, we introduce a simple yet effective image sharpening and denoising technique, coupled with a network that estimates the pixel-wise uncertainty of the predicted surface normal vectors. Identifying such uncertainty can prevent our model from being misled by unreliable surface normal supervisions that hinder the accurate reconstruction of intricate geometries. Experiments on the benchmark datasets show that our method outperforms existing methods in terms of reconstruction quality. Furthermore, the proposed method also generalizes well to real-world indoor scenarios captured by our hand-held mobile phones. Yubin Hu 0001, Matthieu Lin, Yu-Hui Wen, Wang Zhao 0001, Yong-Jin Liu 0001, Wenping Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | Edge-Assisted Collaborative Perception Against Jamming and Interference in Vehicular NetworksabstractCollaborative perception of connected autonomous vehicles (CAVs) that offload the sensing data, such as the feature map extracted from light detection and ranging (LiDAR) point clouds, to an edge device such as the roadside unit (RSU) to detect traffic objects has severe performance degradation due to the offloading latency and packet loss rate (PLR) under jamming and interference. In this paper, we propose an edge-assisted reinforcement learning (RL)-based collaborative perception scheme for CAVs to enhance the accuracy and speed against jamming and interference in LiDAR-based object detection. Based on the spatial confidence score of the feature map, the data size, the channel gains, the received jamming power and interference level, this scheme chooses the critical regions of the feature map, radio channel and transmit power with the hierarchical structure to enhance the learning efficiency. The risk level of the selected policy evaluates the time asynchronization and information loss of the shared feature map using the multi-level risk function based on multiple thresholds of the offloading latency and PLR, with assigning different penalties to mitigate the selection of high-risk policies that degrade perception performance. The upper performance bound in terms of the perception accuracy, latency and utility is provided based on the Stackelberg equilibrium of the game between the jammer and CAVs. Experimental results based on the Robosense RS-LiDAR-16 sensors and the Raspberry Pi to detect 10 vehicles in an$8.5\times 4\times 3.5$m3area show the performance gain with 22.4% higher perception accuracy and 41.3% less latency compared with the benchmark against a smart jammer. Zhiping Lin 0002, Liang Xiao 0003, Zefang Lv, Yunjun Zhu, Yanyong Zhang, Yong-Jin Liu 0001 |
IEEE Trans. Wirel. Commun. | 7 |
| 2024 | O^2-Recon: Completing 3D Reconstruction of Occluded Objects in the Scene with a Pre-trained 2D Diffusion ModelabstractOcclusion is a common issue in 3D reconstruction from RGB-D videos, often blocking the complete reconstruction of objects and presenting an ongoing problem. In this paper, we propose a novel framework, empowered by a 2D diffusion-based in-painting model, to reconstruct complete surfaces for the hidden parts of objects. Specifically, we utilize a pre-trained diffusion model to fill in the hidden areas of 2D images. Then we use these in-painted images to optimize a neural implicit surface representation for each instance for 3D reconstruction. Since creating the in-painting masks needed for this process is tricky, we adopt a human-in-the-loop strategy that involves very little human engagement to generate high-quality masks. Moreover, some parts of objects can be totally hidden because the videos are usually shot from limited perspectives. To ensure recovering these invisible areas, we develop a cascaded network architecture for predicting signed distance field, making use of different frequency bands of positional encoding and maintaining overall smoothness. Besides the commonly used rendering loss, Eikonal loss, and silhouette loss, we adopt a CLIP-based semantic consistency loss to guide the surface from unseen camera angles. Experiments on ScanNet scenes show that our proposed framework achieves state-of-the-art accuracy and completeness in object-level reconstruction from scene-level RGB-D videos. Code: https://github.com/THU-LYJ-Lab/O2-Recon. Yubin Hu 0001, Wang Zhao 0001, Matthieu Lin, Yu-Hui Wen, Ying He 0001, Yong-Jin Liu 0001 |
AAAI | 8 |
| 2024 | SpaceGTN: A Time-Agnostic Graph Transformer Network for Handwritten Diagram Recognition and SegmentationabstractOnline handwriting recognition is pivotal in domains like note-taking, education, healthcare, and office tasks. Existing diagram recognition algorithms mainly rely on the temporal information of strokes, resulting in a decline in recognition performance when dealing with notes that have been modified or have no temporal information. The current datasets are drawn based on templates and cannot reflect the real free-drawing situation. To address these challenges, we present SpaceGTN, a time-agnostic Graph Transformer Network, leveraging spatial integration and removing the need for temporal data. Extensive experiments on multiple datasets have demonstrated that our method consistently outperforms existing methods and achieves state-of-the-art performance. We also propose a pipeline that seamlessly connects offline and online handwritten diagrams. By integrating a stroke restoration technique with SpaceGTN, it enables intelligent editing of previously uneditable offline diagrams at the stroke level. In addition, we have also launched the first online handwritten diagram dataset, OHSD, which is collected using a free-drawing method and comes with modification annotations. Haoxiang Hu, Cangjun Gao, Yaokun Li, Xiaoming Deng 0001, Yukun Lai, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
AAAI | 7 |
| 2024 | Exploring Temporal Feature Correlation for Efficient and Stable Video Semantic SegmentationabstractThis paper tackles the problem of efficient and stable video semantic segmentation. While stability has been under-explored, prevalent work in efficient video semantic segmentation uses the keyframe paradigm. They efficiently process videos by only recomputing the low-level features and reusing high-level features computed at selected keyframes. In addition, the reused features stabilize the predictions across frames, thereby improving video consistency. However, dynamic scenes in the video can easily lead to misalignments between reused and recomputed features, which hampers performance. Moreover, relying on feature reuse to improve prediction consistency is brittle; an erroneous alignment of the features can easily lead to unstable predictions. Therefore, the keyframe paradigm exhibits a dilemma between stability and performance. We address this efficiency and stability challenge using a novel yet simple Temporal Feature Correlation (TFC) module. It uses the cosine similarity between two frames’ low-level features to inform the semantic label’s consistency across frames. Specifically, we selectively reuse label-consistent features across frames through linear interpolation and update others through sparse multi-scale deformable attention. As a result, we no longer directly reuse features to improve stability and thus effectively solve feature misalignment. This work provides a significant step towards efficient and stable video semantic segmentation. On the VSPW dataset, our method significantly improves the prediction consistency of image-based methods while being as fast and accurate. Matthieu Lin, Jenny Sheng, Yubin Hu 0001, Yangguang Li 0001, Andrew Zhao, Gao Huang 0001, Yong-Jin Liu 0001 |
AAAI | 8 |
| 2024 | ExpeL: LLM Agents Are Experiential LearnersabstractThe recent surge in research interest in applying large language models (LLMs) to decision-making tasks has flourished by leveraging the extensive world knowledge embedded in LLMs. While there is a growing demand to tailor LLMs for custom decision-making tasks, finetuning them for specific tasks is resource-intensive and may diminish the model's generalization capabilities. Moreover, state-of-the-art language models like GPT-4 and Claude are primarily accessible through API calls, with their parametric weights remaining proprietary and unavailable to the public. This scenario emphasizes the growing need for new methodologies that allow learning from agent experiences without requiring parametric updates. To address these problems, we introduce the Experiential Learning (ExpeL) agent. Our agent autonomously gathers experiences and extracts knowledge using natural language from a collection of training tasks. At inference, the agent recalls its extracted insights and past experiences to make informed decisions. Our empirical results highlight the robust learning efficacy of the ExpeL agent, indicating a consistent enhancement in its performance as it accumulates experiences. We further explore the emerging capabilities and transfer learning potential of the ExpeL agent through qualitative observations and additional experiments. Andrew Zhao, Daniel Huang 0007, Quentin Xu, Matthieu Lin, Yong-Jin Liu 0001, Gao Huang 0001 |
AAAI | 5 |
| 2024 | Towards More Accurate Diffusion Model Acceleration with a Timestep TunerabstractA diffusion model, which is formulated to produce an image using thousands of denoising steps, usually suffers from a slow inference speed. Existing acceleration algorithms simplify the sampling by skipping most steps yet exhibit considerable performance degradation. By viewing the generation of diffusion models as a discretized integral process, we argue that the quality drop is partly caused by applying an inaccurate integral direction to a timestep interval. To rectify this issue, we propose a timestep tuner that helps find a more accurate integral direction for a particular interval at the minimum cost. Specifically, at each denoising step, we replace the original parameterization by conditioning the network on a new timestep, enforcing the sampling distribution towards the real one. Extensive experiments show that our plug-in design can be trained efficiently and boost the inference performance of various state-of-the-art acceleration methods, especially when there are few denoising steps. For example, when using 10 denoising steps on LSUN Bedroom dataset, we improve the FID of DDIM from 9.65 to 6.07, simply by adopting our method for a more appropriate set of timesteps. Code is available at https://github.com/THU-LYJ-Lab/time-tuner. Mengfei Xia, Yujun Shen, Changsong Lei, Yu Zhou 0076, Deli Zhao, Ran Yi 0002, Wenping Wang 0001, Yong-Jin Liu 0001 |
CVPR | 8 |
| 2024 | NGP-RT: Fusing Multi-level Hash Features with Lightweight Attention for Real-Time Novel View Synthesis
Yubin Hu 0001, Jingwei Huang 0001, Yong-Jin Liu 0001 |
ECCV (22) | 5 |
| 2024 | SMaRt: Improving GANs with Score Matching RegularityabstractGenerative adversarial networks (GANs) usually struggle in learning from highly diverse data, whose underlying manifold is complex. In this work, we revisit the mathematical foundations of GANs, and theoretically reveal that the native adversarial loss for GAN training is insufficient to fix the problem of $\textit{subsets with positive Lebesgue measure of the generated data manifold lying out of the real data manifold}$. Instead, we find that score matching serves as a promising solution to this issue thanks to its capability of persistently pushing the generated data points towards the real data manifold. We thereby propose to improve the optimization of GANs with score matching regularity (SMaRt). Regarding the empirical evidences, we first design a toy example to show that training GANs by the aid of a ground-truth score function can help reproduce the real data distribution more accurately, and then confirm that our approach can consistently boost the synthesis performance of various state-of-the-art GANs on real-world datasets with pre-trained diffusion models acting as the approximate score function. For instance, when training Aurora on the ImageNet $64\times64$ dataset, we manage to improve FID from 8.87 to 7.11, on par with the performance of one-step consistency model. Code is available at https://github.com/thuxmf/SMaRt. Mengfei Xia, Yujun Shen, Ceyuan Yang, Ran Yi 0002, Wenping Wang 0001, Yong-Jin Liu 0001 |
ICML | 6 |
| 2024 | Generalizable Thermal-based Depth Estimation via Pre-trained Visual Foundation ModelabstractDepth estimation is a crucial task in computer vision, applicable to various domains such as 3D reconstruction, robotics, and autonomous driving. In particular, thermal-based depth estimation has unique advantages, including night-time vision. However, the existing depth estimation method remains challenging in robust generalization due to limited data resources and spectral differences between thermal and RGB images. In this paper, we present a self-supervised approach to enhance thermal-based depth estimation by leveraging pre-trained visual models initially designed for RGB data. In detail, we design a novel two-stage training strategy, incorporating Low-rank Adapters and Convolutional Adapters, which not only significantly improves accuracy and robustness but also enables impressive zero-shot generalization capabilities. Our method outperforms existing thermal-based depth estimation models, opening new possibilities for cross-modal applications in computer vision and robotics research. Ruoyu Fan, Wang Zhao 0001, Matthieu Lin, Qi Wang 0079, Yong-Jin Liu 0001, Wenping Wang 0001 |
ICRA | 5 |
| 2024 | MMPI: a Flexible Radiance Field Representation by Multiple Multi-plane Images BlendingabstractThis paper presents a flexible representation of neural radiance fields based on multi-plane images (MPI), for high-quality view synthesis of complex scenes. MPI with Normalized Device Coordinate (NDC) parameterization is widely used in NeRF learning for its simple definition, easy calculation, and powerful ability to represent unbounded scenes. However, existing NeRF works that adopt MPI representation for novel view synthesis can only handle simple forward-facing unbounded scenes (e.g., the scenes in the LLFF dataset), where the input cameras are all observing in similar directions with small relative translations. Hence, extending these MPIbased methods to more complex scenes like large-range or even 360-degree scenes is very challenging. In this paper, we explore the potential of MPI and show that MPI can synthesize high-quality novel views of complex scenes with diverse camera distributions and view directions, which are not only limited to simple forward-facing scenes. Our key idea is to encode the neural radiance field with multiple MPIs facing different directions and blend them with an adaptive blending operation. For each region of the scene, the blending operation gives larger blending weights to those advantaged MPIs with stronger local representation abilities while giving lower weights to those with weaker representation abilities. Such blending operation automatically modulates the multiple MPIs to appropriately represent the diverse local density and color information. Experiments on the KITTI dataset and ScanNet dataset demonstrate that our proposed MMPI synthesizes high-quality images from diverse camera pose distributions and is fast to train, outperforming the previous fast-training NeRF methods for novel view synthesis. Moreover, we show that MMPI can encode extremely long trajectories and produce novel view renderings, demonstrating its potential in applications like autonomous driving. Our demo video is available at https://youtube.com/watch?v=mbNKwN5urC8. Peng Wang 0099, Yubin Hu 0001, Wang Zhao 0001, Ran Yi 0002, Yong-Jin Liu 0001, Wenping Wang 0001 |
ICRA | 6 |
| 2024 | FF-LOGO: Cross-Modality Point Cloud Registration with Feature Filtering and Local to Global OptimizationabstractCross-modality point cloud registration is confronted with significant challenges due to inherent differences in modalities between sensors. To deal with this problem, we propose FF-LOGO: a cross-modality point cloud registration framework with Feature Filtering and LOcal-Global Optimization. The cross-modality feature correlation filtering module extracts geometric transformation-invariant features from cross-modality point clouds and achieves point selection by feature matching. We also introduce a cross-modality optimization process, including a local adaptive key region aggregation module and a global modality consistency fusion optimization module. Experimental results demonstrate that our two-stage optimization significantly improves the registration accuracy of the feature association and selection module. Our method achieves a substantial increase in recall rate compared to the current state-of-the-art methods on the 3DCSR dataset, improving from 40.59% to 75.74%. Our code will be available at https://github.com/wangmohan17/FFLOGO. Nan Ma 0012, Yiheng Han, Yong-Jin Liu 0001 |
ICRA | 4 |
| 2024 | Mobile Robot Oriented Large-Scale Indoor Dataset for Dynamic Scene UnderstandingabstractMost existing robotic datasets capture static scene data and thus are limited in evaluating robots’ dynamic performance. To address this, we present a mobile robot oriented large-scale indoor dataset, denoted as THUD (Tsinghua University Dynamic) robotic dataset, for training and evaluating their dynamic scene understanding algorithms. Specifically, the THUD dataset construction is first detailed, including organization, acquisition, and annotation methods. It comprises both real-world and synthetic data, collected with a real robot platform and a physical simulation platform, respectively. Our current dataset includes 13 larges-scale dynamic scenarios, 90K image frames, 20M 2D/3D bounding boxes of static and dynamic objects, camera poses, and IMU. The dataset is still continuously expanding. Then, the performance of mainstream indoor scene understanding tasks, e.g. 3D object detection, semantic segmentation, and robot relocalization, is evaluated on our THUD dataset. These experiments reveal serious challenges for some robot scene understanding tasks in dynamic scenes. By sharing this dataset, we aim to foster and iterate new mobile robot algorithms quickly for robot actual working dynamic environment, i.e. complex crowded dynamic scenes. Cong Tai, Fang-xing Chen, Wanting Zhang, Tao Zhang 0130, Xueping Liu 0003, Yong-Jin Liu 0001, Long Zeng 0001 |
ICRA | 7 |
| 2024 | SceneDiff: Generative Scene-Level Image Retrieval with Text and Sketch Using Diffusion Models
Ran Zuo, Haoxiang Hu, Xiaoming Deng 0001, Cangjun Gao, Yukun Lai, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
IJCAI | 8 |
| 2024 | MonoPlane: Exploiting Monocular Geometric Cues for Generalizable 3D Plane ReconstructionabstractThis paper presents a generalizable 3D plane detection and reconstruction framework named MonoPlane. Unlike previous robust estimator-based works (which require multiple images or RGB-D input) and learning-based works (which suffer from domain shift), MonoPlane combines the best of two worlds and establishes a plane reconstruction pipeline based on monocular geometric cues, resulting in accurate, robust and scalable 3D plane detection and reconstruction in the wild. Specifically, we first leverage large-scale pre-trained neural networks to obtain the depth and surface normals from a single image. These monocular geometric cues are then incorporated into a proximity-guided RANSAC framework to sequentially fit each plane instance. We exploit effective 3D point proximity and model such proximity via a graph within RANSAC to guide the plane fitting from noisy monocular depths, followed by image-level multi-plane joint optimization to improve the consistency among all plane instances. We further design a simple but effective pipeline to extend this single-view solution to sparse-view 3D plane reconstruction. Extensive experiments on a list of datasets demonstrate our superior zero-shot generalizability over baselines, achieving state-of-the-art plane reconstruction performance in a transferring setting. Our code is available at https://github.com/thuzhaowang/MonoPlane. Wang Zhao 0001, Yishu Li, Sili Chen, Sharon X. Huang, Yong-Jin Liu 0001, Hengkai Guo |
IROS | 7 |
| 2024 | ECAvatar: 3D Avatar Facial Animation with Controllable Identity and Emotion
Minjing Yu, Delong Pang, Ziwen Kang, Zhiyao Sun, Tian Lv, Jenny Sheng, Ran Yi 0002, Yu-Hui Wen, Yong-Jin Liu 0001 |
ACM Multimedia | 9 |
| 2024 | VisHanfu: An Interactive System for the Promotion of Hanfu Knowledge via Cross-Shaped Flat Structure
Minjing Yu, Lingzhi Zeng, Xinxin Du, Jenny Sheng, Qiantian Liao, Yong-Jin Liu 0001 |
ACM Multimedia | 6 |
| 2024 | SS3DM: Benchmarking Street-View Surface Reconstruction with a Synthetic 3D Mesh DatasetabstractReconstructing accurate 3D surfaces for street-view scenarios is crucial for applications such as digital entertainment and autonomous driving simulation. However, existing street-view datasets, including KITTI, Waymo, and nuScenes, only offer noisy LiDAR points as ground-truth data for geometric evaluation of reconstructed surfaces. These geometric ground-truths often lack the necessary precision to evaluate surface positions and do not provide data for assessing surface normals. To overcome these challenges, we introduce the SS3DM dataset, comprising precise \textbf{S}ynthetic \textbf{S}treet-view \textbf{3D} \textbf{M}esh models exported from the CARLA simulator. These mesh models facilitate accurate position evaluation and include normal vectors for evaluating surface normal. To simulate the input data in realistic driving scenarios for 3D reconstruction, we virtually drive a vehicle equipped with six RGB cameras and five LiDAR sensors in diverse outdoor scenes. Leveraging this dataset, we establish a benchmark for state-of-the-art surface reconstruction methods, providing a comprehensive evaluation of the associated challenges. For more information, visit our homepage at https://ss3dm.top. Yubin Hu 0001, Kairui Wen, Yong-Jin Liu 0001 |
NeurIPS | 5 |
| 2024 | AlphaTablets: A Generic Plane Representation for 3D Planar Reconstruction from Monocular VideosabstractWe introduce AlphaTablets, a novel and generic representation of 3D planes that features continuous 3D surface and precise boundary delineation. By representing 3D planes as rectangles with alpha channels, AlphaTablets combine the advantages of current 2D and 3D plane representations, enabling accurate, consistent and flexible modeling of 3D planes. We derive differentiable rasterization on top of AlphaTablets to efficiently render 3D planes into images, and propose a novel bottom-up pipeline for 3D planar reconstruction from monocular videos. Starting with 2D superpixels and geometric cues from pre-trained models, we initialize 3D planes as AlphaTablets and optimize them via differentiable rendering. An effective merging scheme is introduced to facilitate the growth and refinement of AlphaTablets. Through iterative optimization and merging, we reconstruct complete and accurate 3D planes with solid surfaces and clear boundaries. Extensive experiments on the ScanNet dataset demonstrate state-of-the-art performance in 3D planar reconstruction, underscoring the great potential of AlphaTablets as a generic 3D plane representation for various applications. Wang Zhao 0001, Shaohui Liu, Yubin Hu 0001, Yushi Bai, Yu-Hui Wen, Yong-Jin Liu 0001 |
NeurIPS | 7 |
| 2024 | Text-image conditioned diffusion for consistent text-to-3D generation
Yushi Bai, Matthieu Lin, Jenny Sheng, Yubin Hu 0001, Qi Wang 0079, Yu-Hui Wen, Yong-Jin Liu 0001 |
Comput. Aided Geom. Des. | 8 |
| 2024 | Automatic tooth arrangement with joint features of point and mesh representations via diffusion probabilistic models
Changsong Lei, Mengfei Xia, Shaofeng Wang, Yaqian Liang, Ran Yi 0002, Yu-Hui Wen, Yong-Jin Liu 0001 |
Comput. Aided Geom. Des. | 7 |
| 2024 | Gaussian in the Dark: Real-Time View Synthesis From Inconsistent Dark Images Using Gaussian SplattingabstractAbstract 3D Gaussian Splatting has recently emerged as a powerful representation that can synthesize remarkable novel views using consistent multi‐view images as input. However, we notice that images captured in dark environments where the scenes are not fully illuminated can exhibit considerable brightness variations and multi‐view inconsistency, which poses great challenges to 3D Gaussian Splatting and severely degrades its performance. To tackle this problem, we propose Gaussian‐DK. Observing that inconsistencies are mainly caused by camera imaging, we represent a consistent radiance field of the physical world using a set of anisotropic 3D Gaussians, and design a camera response module to compensate for multi‐view inconsistencies. We also introduce a step‐based gradient scaling strategy to constrain Gaussians near the camera, which turn out to be floaters, from splitting and cloning. Experiments on our proposed benchmark dataset demonstrate that Gaussian‐DK produces high‐quality renderings without ghosting and floater artifacts and significantly outperforms existing methods. Furthermore, we can also synthesize light‐up images by controlling exposure levels that clearly show details in shadow areas. Zhen-Hui Dong, Yubin Hu 0001, Yu-Hui Wen, Yong-Jin Liu 0001 |
Comput. Graph. Forum | 5 |
| 2024 | A Diffusion Model Translator for Efficient Image-to-Image TranslationabstractApplying diffusion models to image-to-image translation (I2I) has recently received increasing attention due to its practical applications. Previous attempts inject information from the source image into each denoising step for an iterative refinement, thus resulting in a time-consuming implementation. We propose an efficient method that equips a diffusion model with a lightweight translator, dubbed a Diffusion Model Translator (DMT), to accomplish I2I. Specifically, we first offer theoretical justification that in employing the pioneering DDPM work for the I2I task, it is both feasible and sufficient to transfer the distribution from one domain to another only at some intermediate step. We further observe that the translation performance highly depends on the chosen timestep for domain transfer, and therefore propose a practical strategy to automatically select an appropriate timestep for a given task. We evaluate our approach on a range of I2I applications, including image stylization, image colorization, segmentation to image, and sketch to image, to validate its efficacy and general utility. The comparisons show that our DMT surpasses existing methods in both quality and efficiency. Code is available at https://github.com/THU-LYJ-Lab/dmt. Mengfei Xia, Yu Zhou 0076, Ran Yi 0002, Yong-Jin Liu 0001, Wenping Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | FEditNet++: Few-Shot Editing of Latent Semantics in GAN Spaces With Correlated Attribute DisentanglementabstractGenerative Adversarial Networks have achieved significant advancements in generating and editing high-resolution images. However, most methods suffer from either requiring extensive labeled datasets or strong prior knowledge. It is also challenging for them to disentangle correlated attributes with few-shot data. In this paper, we propose FEditNet++, a GAN-based approach to explore latent semantics. It aims to enable attribute editing with limited labeled data and disentangle the correlated attributes. We propose a layer-wise feature contrastive objective, which takes into consideration content consistency and facilitates the invariance of the unrelated attributes before and after editing. Furthermore, we harness the knowledge from the pretrained discriminative model to prevent overfitting. In particular, to solve the entanglement problem between the correlated attributes from data and semantic latent correlation, we extend our model to jointly optimize multiple attributes and propose a novel decoupling loss and cross-assessment loss to disentangle them from both latent and image space. We further propose a novel-attribute disentanglement strategy to enable editing of novel attributes with unknown entanglements. Finally, we extend our model to accurately edit the fine-grained attributes. Qualitative and quantitative assessments demonstrate that our method outperforms state-of-the-art approaches across various datasets, including CelebA-HQ, RaFD, Danbooru2018 and LSUN Church. Ran Yi 0002, Mengfei Xia, Yizhe Tang, Yong-Jin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Focus on Cooperation: A Face-to-Face VR Serious Game for Relationship EnhancementabstractExploring effective approaches to enhance face-to-face interactions and interpersonal relationships is an important topic in the applications of affective computing. According to the co-actualization model, we propose a face-to-face co-participation serious game for relationship enhancement, with a focus on battling COVID-19. Moreover, a prototype system is developed using an immersive virtual environment and a low-cost brain-computer interface. Through this system, a dynamic flow experience enhancement tool is utilized to involve partners in the cooperative task. To evaluate the system performance, two studies are conducted with schoolmates as participants. Study 1 compares the cooperative and competitive modes, and demonstrates that the former elicited higher level of decision-making challenge and affections, which are beneficial for forming relationships. Study 2 further examines the effect of the dynamic flow enhancement tool in the cooperative task and the results show its effectiveness in promoting flow experience, perceived closeness, and intimacy in relationships. Given this short-term participation, participants felt a greater sense of closeness and intimacy than they had before the test. In conclusion, our proposed system is effective in enhancing schoolmate relationships. Yulong Bian, Chao Zhou 0012, Yang Zhang 0116, Juan Liu 0008, Jenny Sheng, Yong-Jin Liu 0001 |
IEEE Trans. Affect. Comput. | 6 |
| 2024 | Emotion Recognition From Few-Channel EEG Signals by Integrating Deep Feature Aggregation and Transfer LearningabstractElectroencephalogram (EEG) signals have been widely studied in human emotion recognition. The majority of existing EEG emotion recognition algorithms utilize dozens or hundreds of electrodes covering the whole scalp region (denoted as full-channel EEG devices in this paper). Nowadays, more and more portable and miniature EEG devices with only a few electrodes (denoted as few-channel EEG devices in this paper) are emerging. However, emotion recognition from few-channel EEG data is challenging because the device can only capture EEG signals from a portion of the brain area. Moreover, existing full-channel algorithms cannot be directly adapted to few-channel EEG signals due to the significant inter-variation between full-channel and few-channel EEG devices. To address these challenges, we propose a novel few-channel EEG emotion recognition framework from the perspective of knowledge transfer. We leverage full-channel EEG signals to provide supplementary information for few-channel signals via a transfer learning-based model CD-EmotionNet, which consists of a base emotion model for efficient emotional feature extraction and a cross-device transfer learning strategy. This strategy helps to enhance emotion recognition performance on few-channel EEG data by utilizing knowledge learned from full-channel EEG data. To evaluate our cross-device EEG emotion transfer learning framework, we construct an emotion dataset containing paired 18-channel and 5-channel EEG signals from 25 subjects, as well as 5-channel EEG signals from 13 other subjects. Extensive experiments show that our framework outperforms state-of-the-art EEG emotion recognition methods by a large margin. Fang Liu 0035, Yezhi Shu, Niqi Liu, Jenny Sheng, Xiaoan Wang, Yong-Jin Liu 0001 |
IEEE Trans. Affect. Comput. | 8 |
| 2024 | Emotion Dictionary Learning With Modality Attentions for Mixed Emotion ExplorationabstractMulti-modal emotion analysis, as an important direction in affective computing, has attracted increasing attention in recent years. Most existing multi-modal emotion recognition studies are targeted at a classification task that aims to assign a specific emotion category to a combination of several heterogeneous input data, including multimedia signals and physiological signals. Compared to single-class emotion recognition, a growing number of recent psychological evidence suggests that different discrete emotions may co-exist at the same time, which promotes the development of mixed-emotion recognition to identify a mixture of basic emotions. Although most current studies treat it as a multi-label classification task, in this work, we focus on a challenging situation where both positive and negative emotions are presented simultaneously, and propose a multi-modal mixed emotion recognition framework, namely EmotionDict. The key characteristics of our EmotionDict include the following. (1) Inspired by the psychological evidence that such a mixed state can be represented by combinations of basic emotions, we address mixed emotion recognition as a label distribution learning task. An emotion dictionary has been designed to disentangle the mixed emotion representations into a weighted sum of a set of basic emotion elements in a shared latent space and their corresponding weights. (2) While many existing emotion distribution studies are built on a single type of multimedia signal (such as text, image, audio, and video), we incorporate physiological and overt behavioral multi-modal signals, including electroencephalogram (EEG), peripheral physiological signals, and facial videos, which directly display the subjective emotions. These modalities have diverse characteristics given that they are related to the central or peripheral nervous system, and the motor cortex. (3) We further design auxiliary tasks to learn modality attentions for modality integration. Experiments on two datasets show that our method outperforms existing state-of-the-art approaches on mixed-emotion recognition. Fang Liu 0035, Yezhi Shu, Fei Yan 0002, Yong-Jin Liu 0001 |
IEEE Trans. Affect. Comput. | 6 |
| 2024 | Continuously Controllable Facial Expression Editing in Talking Face VideosabstractRecently audio-driven talking face video generation has attracted considerable attention. However, very few researches address the issue of emotional editing of these talking face videos with continuously controllable expressions, which is a strong demand in the industry. The challenge is that speech-related expressions and emotion-related expressions are often highly coupled. Meanwhile, traditional image-to-image translation methods cannot work well in our application due to the coupling of expressions with other attributes such as poses, i.e., translating the expression of the character in each frame may simultaneously change the head pose due to the bias of the training data distribution. In this paper, we propose a high-quality facial expression editing method for talking face videos, allowing the user to control the target emotion in the edited video continuously. We present a new perspective for this task as a special case of motion information editing, where we use a 3DMM to capture major facial movements and an associated texture map modeled by a StyleGAN to capture appearance details. Both representations (3DMM and texture map) contain emotional information and can be continuously modified by neural networks and easily smoothed by averaging in coefficient/latent spaces, making our method simple yet effective. We also introduce a mouth shape preservation loss to control the trade-off between lip synchronization and the degree of exaggeration of the edited expression. Extensive experiments and a user study show that our method achieves state-of-the-art performance across various evaluation criteria. Zhiyao Sun, Yu-Hui Wen, Tian Lv, Yanan Sun 0006, Yaoyuan Wang, Yong-Jin Liu 0001 |
IEEE Trans. Affect. Comput. | 7 |
| 2024 | Reinforcement Learning Based Energy-Efficient Fast Routing for FANETsabstractReinforcement learning (RL) based flying ad-hoc network (FANET) routing enables unmanned aerial vehicles (UAVs) to choose the next-hop to increase the packet delivery ratio, but the routing latency and energy consumption have to be further reduced over inaccurate feedback for large-scale networks. In this paper, we propose an RL based energy-efficient fast routing for each UAV to choose the forwarding decision and the power. Based on the state consisting of the battery level, channel conditions and forwarding decisions of the one-hop neighbors, the routing policy is chosen to enhance the utility as the weighted sum of the delivery success indicator, the latency and the energy consumption. The number of the latency violations and the learning parameters shared among the one-hop neighbors are exploited in the update of the routing policy distribution following the latency constraint with the reduced energy consumption. The deep neural networks address the state quantization error of the latency and the channel gain for UAVs with high mobility under large-scale networks. The performance bound regarding the end-to-end latency and the energy consumption is derived in terms of network topology and channel gain based on the packet forwarding game. The performance gain over the benchmark is provided via both simulation and experimental results. Jieling Li, Liang Xiao 0003, Xuchen Qi, Zefang Lv, Qiaoxin Chen, Yong-Jin Liu 0001 |
IEEE Trans. Commun. | 6 |
| 2024 | MFDAN: Multi-Level Flow-Driven Attention Network for Micro-Expression RecognitionabstractFacial expressions are an essential part of human emotional communication, and micro-expressions (MEs), as transient and imperceptible non-verbal signals, can potentially reveal real human emotions. However, subtle motion variations, limited and unbalanced samples make micro-expression recognition (MER) challenging. In this paper, we design a novel dual-branch learning framework of multi-level flow-driven attention for micro-expression recognition (MFDAN), which innovatively integrates optical flow prior to guide the attention learning in the image encoding branch, enabling the model to focus on the most discriminative facial regions for subtle motion patterns. Firstly, we extract optical flow information by an optical flow encoding module. Then, in the image coding module, we construct a Transformer structure containing an optical flow-driven attention mechanism, which can effectively locate the interest region of micro-expressions in the image according to the position information of optical flow to capture more sensitive and fine-grained micro-expressions. By interoperating prior knowledge with data learning, and introducing the Dropkey operation and Focal Loss, our method can handle subtle micro-expression features on small imbalanced datasets. Through extensive experiments on three independent datasets and a composite database, including SMIC-HS, SAMM, and CASME II, robust leave-one-subject-out (LOSO) evaluation results show that our method outperforms state-of-the-art methods especially on the composite database. Junli Zhao, Ran Yi 0002, Minjing Yu, Fuqing Duan, Zhenkuan Pan 0001, Yong-Jin Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | SD-FSOD: Self-Distillation Paradigm via Distribution Calibration for Few-Shot Object DetectionabstractFew-shot object detection (FSOD) aims to detect novel targets with only a few instances of the associated samples. Although combinations of distillation techniques and meta-learning paradigms have been acknowledged as the primary strategies for FSOD tasks, the existing distillation methods exhibit inherent biases and sensitivity to novel class variability. A critical hurdle for FSOD distillation is the difficulty in ensuring appropriate knowledge learned from the teacher model during the fine-tuning stage. Furthermore, coarse distillation procedures risk misalignment between the learned and actual distributions. This misalignment could potentially negate the benefits of positive cases and impede the detector’s evolution. To address these deficiencies, we propose a novel self-distillation paradigm exclusively for the fine-tuning stage (SD-FSOD). Our methods integrate a Distribution Prototype Extractor (DPE) and Self-Distillation Memory (SDM), promoting feature distribution consistency during distillation. In detail, the DPE module reliably initializes the weights of the detector, ensuring a robust class distribution for the distillation process. Meanwhile, the SDM module utilizes decoupling techniques to divide the distillation tasks into two sub-task branches, allowing the student model to independently learn and share precise features through isolated distillation processes. The synergistic integration of feature calibration techniques and the continuous self-distillation paradigm distinctly enhances the fine-tuning process, which shows the superiority of the FSOD self-distillation methodologies. The extensive experiments on the PASCAL VOC and MS COCO datasets demonstrate that our proposed approach produces significant improvements and achieves state-of-the-art (SOTA) performance. Qi Wang 0079, Kailin Xie, Liang Lei, Matthieu Lin, Tian Lv, Yong-Jin Liu 0001, Jiebo Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | CDKM: Common and Distinct Knowledge Mining Network With Content Interaction for Dense CaptioningabstractThe dense captioning task aims at detecting multiple salient regions of an image and describing them separately in natural language. Although significant advancements in the field of dense captioning have been made, there are still some limitations to existing methods in recent years. On the one hand, most dense captioning methods lack strong target detection capabilities and struggle to cover all relevant content when dealing with target-intensive images. On the other hand, current transformer-based methods are powerful but neglect the acquisition and utilization of contextual information, hindering the visual understanding of local areas. To address these issues, we propose a common and distinct knowledge-mining network with content interaction for the task of dense captioning. Our network has a knowledge mining mechanism that improves the detection of salient targets by capturing common and distinct knowledge from multi-scale features. We further propose a content interaction module that combines region features into a unique context based on their correlation. Our experiments on various benchmarks have shown that the proposed method outperforms the current state-of-the-art methods. Hongyu Deng, Yushan Xie, Qi Wang 0079, Weijian Ruan, Wu Liu 0005, Yong-Jin Liu 0001 |
IEEE Trans. Multim. | 7 |
| 2024 | DiffPoseTalk: Speech-Driven Stylistic 3D Facial Animation and Head Pose Generation via Diffusion ModelsabstractThe generation of stylistic 3D facial animations driven by speech presents a significant challenge as it requires learning a many-to-many mapping between speech, style, and the corresponding natural facial motion. However, existing methods either employ a deterministic model for speech-to-motion mapping or encode the style using a one-hot encoding scheme. Notably, the one-hot encoding approach fails to capture the complexity of the style and thus limits generalization ability. In this paper, we propose DiffPoseTalk, a generative framework based on the diffusion model combined with a style encoder that extracts style embeddings from short reference videos. During inference, we employ classifier-free guidance to guide the generation process based on the speech and style. In particular, our style includes the generation of head poses, thereby enhancing user perception. Additionally, we address the shortage of scanned 3D talking face data by training our model on reconstructed 3DMM parameters from a high-quality, in-the-wild audio-visual dataset. Extensive experiments and user study demonstrate that our approach outperforms state-of-the-art methods. The code and dataset are at https://diffposetalk.github.io. Zhiyao Sun, Tian Lv, Matthieu Lin, Jenny Sheng, Yu-Hui Wen, Minjing Yu, Yong-Jin Liu 0001 |
ACM Trans. Graph. | 8 |
| 2024 | PVP-Recon: Progressive View Planning via Warping Consistency for Sparse-View Surface ReconstructionabstractNeural implicit representations have revolutionized dense multi-view surface reconstruction, yet their performance significantly diminishes with sparse input views. A few pioneering works have sought to tackle this challenge by leveraging additional geometric priors or multi-scene generalizability. However, they are still hindered by the imperfect choice of input views, using images under empirically determined viewpoints. We propose PVP-Recon , a novel and effective sparse-view surface reconstruction method that progressively plans the next best views to form an optimal set of sparse viewpoints for image capturing. PVP-Recon starts initial surface reconstruction with as few as 3 views and progressively adds new views which are determined based on a novel warping score that reflects the information gain of each newly added view. This progressive view planning progress is interleaved with a neural SDF-based reconstruction module that utilizes multi-resolution hash features, enhanced by a progressive training scheme and a directional Hessian loss. Quantitative and qualitative experiments on three benchmark datasets show that our system achieves high-quality reconstruction with a constrained input budget and outperforms existing baselines. Matthieu Lin, Jenny Sheng, Ruoyu Fan, Yiheng Han, Yubin Hu 0001, Ran Yi 0002, Yu-Hui Wen, Yong-Jin Liu 0001, Wenping Wang 0001 |
ACM Trans. Graph. | 10 |
| 2024 | SpeechMirror: A Multimodal Visual Analytics System for Personalized Reflection of Online Public Speaking EffectivenessabstractAs communications are increasingly taking place virtually, the ability to present well online is becoming an indispensable skill. Online speakers are facing unique challenges in engaging with remote audiences. However, there has been a lack of evidence-based analytical systems for people to comprehensively evaluate online speeches and further discover possibilities for improvement. This paper introduces SpeechMirror, a visual analytics system facilitating reflection on a speech based on insights from a collection of online speeches. The system estimates the impact of different speech techniques on effectiveness and applies them to a speech to give users awareness of the performance of speech techniques. A similarity recommendation approach based on speech factors or script content supports guided exploration to expand knowledge of presentation evidence and accelerate the discovery of speech delivery possibilities. SpeechMirror provides intuitive visualizations and interactions for users to understand speech factors. Among them, SpeechTwin, a novel multimodal visual summary of speech, supports rapid understanding of critical speech factors and comparison of different speech samples, and SpeechPlayer augments the speech video by integrating visualization of the speaker's body language with interaction, for focused analysis. The system utilizes visualizations suited to the distinct nature of different speech factors for user comprehension. The proposed system and visualization techniques were evaluated with domain experts and amateurs, demonstrating usability for users with low visualization literacy and its efficacy in assisting users to develop insights for potential improvement. Kevin T. Maher, Xiaoming Deng 0001, Yukun Lai, CuiXia Ma, Sheng Feng Qin, Yong-Jin Liu 0001, Hongan Wang |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2024 | MSL-Net: Sharp Feature Detection Network for 3D Point CloudsabstractAs a significant geometric feature of 3D point clouds, sharp features play an important role in shape analysis, 3D reconstruction, registration, localization, etc. Current sharp feature detection methods are still sensitive to the quality of the input point cloud, and the detection performance is affected by random noisy points and non-uniform densities. In this paper, using the prior knowledge of geometric features, we propose a Multi-scale Laplace Network (MSL-Net), a new deep-learning-based method based on an intrinsic neighbor shape descriptor, to detect sharp features from 3D point clouds. First, we establish a discrete intrinsic neighborhood of the point cloud based on the Laplacian graph, which reduces the error of local implicit surface estimation. Then, we design a new intrinsic shape descriptor based on the intrinsic neighborhood, combined with enhanced normal extraction and cosine-based field estimation function. Finally, we present the backbone of MSL-Net based on the intrinsic shape descriptor. Benefiting from the intrinsic neighborhood and shape descriptor, our MSL-Net has simple architecture and is capable of establishing accurate feature prediction that satisfies the manifold distribution while avoiding complex intrinsic metric calculations. Extensive experimental results demonstrate that with the multi-scale structure, MSL-Net has a strong analytical ability for local perturbations of point clouds. Compared with state-of-the-art methods, our MSL-Net is more robust and accurate. Xianhe Jiao, Chenlei Lv, Ran Yi 0002, Junli Zhao, Zhenkuan Pan 0001, Zhongke Wu, Yong-Jin Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2024 | Complete 3D Relationships Extraction Modality Alignment Network for 3D Dense Captioningabstract3D dense captioning aims to semantically describe each object detected in a 3D scene, which plays a significant role in 3D scene understanding. Previous works lack a complete definition of 3D spatial relationships and the directly integrate visual and language modalities, thus ignoring the discrepancies between the two modalities. To address these issues, we propose a novel complete 3D relationship extraction modality alignment network, which consists of three steps: 3D object detection, complete 3D relationships extraction, and modality alignment caption. To comprehensively capture the 3D spatial relationship features, we define a complete set of 3D spatial relationships, including the local spatial relationship between objects and the global spatial relationship between each object and the entire scene. To this end, we propose a complete 3D relationships extraction module based on message passing and self-attention to mine multi-scale spatial relationship features and inspect the transformation to obtain features in different views. In addition, we propose the modality alignment caption module to fuse multi-scale relationship features and generate descriptions to bridge the semantic gap from the visual space to the language space with the prior information in the word embedding, and help generate improved descriptions for the 3D scene. Extensive experiments demonstrate that the proposed model outperforms the state-of-the-art methods on the ScanRefer and Nr3D datasets. Aihua Mao, Wanxin Chen, Ran Yi 0002, Yong-Jin Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | Keyframe Control of Music-Driven 3D Dance GenerationabstractFor 3D animators, choreography with artificial intelligence has attracted more attention recently. However, most existing deep learning methods mainly rely on music for dance generation and lack sufficient control over generated dance motions. To address this issue, we introduce the idea of keyframe interpolation for music-driven dance generation and present a novel transition generation technique for choreography. Specifically, this technique synthesizes visually diverse and plausible dance motions by using normalizing flows to learn the probability distribution of dance motions conditioned on a piece of music and a sparse set of key poses. Thus, the generated dance motions respect both the input musical beats and the key poses. To achieve a robust transition of varying lengths between the key poses, we introduce a time embedding at each timestep as an additional condition. Extensive experiments show that our model generates more realistic, diverse, and beat-matching dance motions than the compared state-of-the-art methods, both qualitatively and quantitatively. Our experimental results demonstrate the superiority of the keyframe-based control for improving the diversity of the generated dance motions. Yu-Hui Wen, Xiao Liu 0040, Yong-Jin Liu 0001, Lin Gao 0004, Hongbo Fu 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | Efficient Communications in Multi-Agent Reinforcement Learning for Mobile ApplicationsabstractThe environment observations and learning experiences shared by the cooperative learning agents accelerate multi-agent reinforcement learning (MARL) with partial observations for mobile applications but the performance degrades due to the redundant and outdated observations under severe channel fading in wireless networks. In this paper, we propose an efficient communication scheme in MARL for mobile applications that enables each learning agent to optimize the cooperative agents and the learning parameters to integrate the shared information. The cooperative agents are chosen according to the learning environment observations, the channel states, and the task similarity with neighboring agents. The learning parameters are chosen based on the attention mechanism that exploits the correlation with the local observation to enhance the agent receptive field for efficient policy exploration. Neural networks with weights updated based on the learning factors determined by the task similarity are designed to further improve the learning efficiency. The performance bounds including the information gain from the learning agent cooperation, the communication cost and the utility are provided based on the Nash equilibrium of the cooperative MARL communication game. The proposed scheme is implemented in the anti-jamming video transmission of the unmanned aerial vehicle swarms to optimize the transmit channel and power and experimental results verify the performance gain over the benchmark. Zefang Lv, Liang Xiao 0003, Yousong Du, Yunjun Zhu, Shuai Han 0002, Yong-Jin Liu 0001 |
IEEE Trans. Wirel. Commun. | 6 |
| 2023 | DarkFeat: Noise-Robust Feature Detector and Descriptor for Extremely Low-Light RAW ImagesabstractLow-light visual perception, such as SLAM or SfM at night, has received increasing attention, in which keypoint detection and local feature description play an important role. Both handcraft designs and machine learning methods have been widely studied for local feature detection and description, however, the performance of existing methods degrades in the extreme low-light scenarios in a certain degree, due to the low signal-to-noise ratio in images. To address this challenge, images in RAW format that retain more raw sensing information have been considered in recent works with a denoise-then-detect scheme. However, existing denoising methods are still insufficient for RAW images and heavily time-consuming, which limits the practical applications of such scheme. In this paper, we propose DarkFeat, a deep learning model which directly detects and describes local features from extreme low-light RAW images in an end-to-end manner. A novel noise robustness map and selective suppression constraints are proposed to effectively mitigate the influence of noise and extract more reliable keypoints. Furthermore, a customized pipeline of synthesizing dataset containing low-light RAW image matching pairs is proposed to extend end-to-end training. Experimental results show that DarkFeat achieves state-of-the-art performance on both indoor and outdoor parts of the challenging MID benchmark, outperforms the denoise-then-detect methods and significantly reduces computational costs up to 70%. Code is available at https://github.com/THU-LYJ-Lab/DarkFeat. Yubin Hu 0001, Wang Zhao 0001, Jisheng Li, Yong-Jin Liu 0001, Yuxing Han 0001, Jiangtao Wen |
AAAI | 5 |
| 2023 | FEditNet: Few-Shot Editing of Latent Semantics in GAN SpacesabstractGenerative Adversarial networks (GANs) have demonstrated their powerful capability of synthesizing high-resolution images, and great efforts have been made to interpret the semantics in the latent spaces of GANs. However, existing works still have the following limitations: (1) the majority of works rely on either pretrained attribute predictors or large-scale labeled datasets, which are difficult to collect in most cases, and (2) some other methods are only suitable for restricted cases, such as focusing on interpretation of human facial images using prior facial semantics. In this paper, we propose a GAN-based method called FEditNet, aiming to discover latent semantics using very few labeled data without any pretrained predictors or prior knowledge. Specifically, we reuse the knowledge from the pretrained GANs, and by doing so, avoid overfitting during the few-shot training of FEditNet. Moreover, our layer-wise objectives which take content consistency into account also ensure the disentanglement between attributes. Qualitative and quantitative results demonstrate that our method outperforms the state-of-the-art methods on various datasets. The code is available at https://github.com/THU-LYJ-Lab/FEditNet. Mengfei Xia, Yezhi Shu, Yuji Wang, Yukun Lai, Qiang Li 0024, Pengfei Wan 0001, Zhongyuan Wang 0006, Yong-Jin Liu 0001 |
AAAI | 8 |
| 2023 | An Easy-to-Build Modular Robot Implementation of Chain-Based Physical Transformation for STEM Education
Minjing Yu, Jeffrey Too Chuan Tan, Yong-Jin Liu 0001 |
CAD/Graphics | 4 |
| 2023 | Efficient Semantic Segmentation by Altering Resolutions for Compressed VideosabstractVideo semantic segmentation (VSS) is a computationally expensive task due to the per-frame prediction for videos of high frame rates. In recent work, compact models or adaptive network strategies have been proposed for efficient VSS. However, they did not consider a crucial factor that affects the computational cost from the input side: the input resolution. In this paper, we propose an altering resolution framework called AR-Seg for compressed videos to achieve efficient VSS. AR-Seg aims to reduce the computational cost by using low resolution for non-keyframes. To prevent the performance degradation caused by downsampling, we design a Cross Resolution Feature Fusion (CR-eFF) module, and supervise it with a novel Feature Similarity Training (FST) strategy. Specifically, CReFF first makes use of motion vectors stored in a compressed video to warp features from high-resolution keyframes to low-resolution non-keyframes for better spatial alignment, and then selectively aggregates the warped features with local attention mechanism. Furthermore, the proposed FST supervises the aggregated features with high-resolution features through an explicit similarity loss and an implicit constraint from the shared decoding layer. Extensive experiments on CamVid and Cityscapes show that AR-Seg achieves state-of-the-art performance and is compatible with different segmentation backbones. On CamVid, AR-Seg saves 67% computational cost (measured in GFLOPs) with the PSPNet18 back-bone while maintaining high segmentation accuracy. Code: https://github.com/THU-LYJ-Lab/AR-Seg. Yubin Hu 0001, Yanghao Li, Jisheng Li, Yuxing Han 0001, Jiangtao Wen, Yong-Jin Liu 0001 |
CVPR | 7 |
| 2023 | Invertible Residual Neural Networks with Conditional Injector and Interpolator for Point Cloud UpsamplingabstractPoint clouds obtained by LiDAR and other sensors are usually sparse and irregular. Low-quality point clouds have serious influence on the final performance of downstream tasks. Recently, a point cloud upsampling network with normalizing flows has been proposed to address this problem. However, the network heavily relies on designing specialized architectures to achieve invertibility. In this paper, we propose a novel invertible residual neural network for point cloud upsampling, called PU-INN, which allows unconstrained architectures to learn more expressive feature transformations. Then, we propose a conditional injector to improve nonlinear transformation ability of the neural network while guaranteeing invertibility. Furthermore, a lightweight interpolator is proposed based on semantic similarity distance in the latent space, which can intuitively reflect the interpolation changes in Euclidean space. Qualitative and quantitative results show that our method outperforms the state-of-the-art works in terms of distribution uniformity, proximity-to-surface accuracy, 3D reconstruction quality, and computation efficiency. Aihua Mao, Yaqi Duan, Yu-Hui Wen, Zihui Du, Hongmin Cai, Yong-Jin Liu 0001 |
IJCAI | 6 |
| 2023 | 4D facial analysis: A survey of datasets, algorithms and applications
Yong-Jin Liu 0001, Baodong Wang, Lin Gao 0004, Junli Zhao, Ran Yi 0002, Minjing Yu, Zhenkuan Pan 0001, Xianfeng Gu |
Comput. Graph. | 1 |
| 2023 | Generation of virtual digital human for customer service industry
Yanan Sun 0006, Zhiyao Sun, Yu-Hui Wen, Tian Lv, Minjing Yu, Ran Yi 0002, Lin Gao 0004, Yong-Jin Liu 0001 |
Comput. Graph. | 9 |
| 2023 | Motif-GCNs With Local and Non-Local Temporal Blocks for Skeleton-Based Action RecognitionabstractRecent works have achieved remarkable performance for action recognition with human skeletal data by utilizing graph convolutional models. Existing models mainly focus on developing graph convolutional operations to encode structural properties of a skeletal graph, whose topology is manually predefined and fixed over all action samples. Some recent works further take sample-dependent relationships among joints into consideration. However, the complex relationships between arbitrary pairwise joints are difficult to learn and the temporal features between frames are not fully exploited by simply using traditional convolutions with small local kernels. In this paper, we propose a motif-based graph convolution method, which makes use of sample-dependent latent relations among non-physically connected joints to impose a high-order locality and assigns different semantic roles to physical neighbors of a joint to encode hierarchical structures. Furthermore, we propose a sparsity-promoting loss function to learn a sparse motif adjacency matrix for latent dependencies in non-physical connections. For extracting effective temporal information, we propose an efficient local temporal block. It adopts partial dense connections to reuse temporal features in local time windows, and enrich a variety of information flow by gradient combination. In addition, we introduce a non-local temporal block to capture global dependencies among frames. Our model can capture local and non-local relationships both spatially and temporally, by integrating the local and non-local temporal blocks into the sparse motif-based graph convolutional networks (SMotif-GCNs). Comprehensive experiments on four large-scale datasets show that our model outperforms the state-of-the-art methods. Our code is publicly available at https://github.com/wenyh1616/SAMotif-GCN. Yu-Hui Wen, Lin Gao 0004, Hongbo Fu 0001, Shihong Xia, Yong-Jin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Quality Metric Guided Portrait Line Drawing Generation From Unpaired Training DataabstractFace portrait line drawing is a unique style of art which is highly abstract and expressive. However, due to its high semantic constraints, many existing methods learn to generate portrait drawings using paired training data, which is costly and time-consuming to obtain. In this paper, we propose a novel method to automatically transform face photos to portrait drawings using unpaired training data with two new features; i.e., our method can (1) learn to generate high quality portrait drawings in multiple styles using a single network and (2) generate portrait drawings in a "new style" unseen in the training data. To achieve these benefits, we (1) propose a novel quality metric for portrait drawings which is learned from human perception, and (2) introduce a quality loss to guide the network toward generating better looking portrait drawings. We observe that existing unpaired translation methods such as CycleGAN tend to embed invisible reconstruction information indiscriminately in the whole drawings due to significant information imbalance between the photo and portrait drawing domains, which leads to important facial features missing. To address this problem, we propose a novel asymmetric cycle mapping that enforces the reconstruction information to be visible and only embedded in the selected facial regions. Along with localized discriminators for important facial regions, our method well preserves all important facial features in the generated drawings. Generator dissection further explains that our model learns to incorporate face semantic information during drawing generation. Extensive experiments including a user study show that our model outperforms state-of-the-art methods. Ran Yi 0002, Yong-Jin Liu 0001, Yukun Lai, Paul L. Rosin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | MMPosE: Movie-Induced Multi-Label Positive Emotion Classification Through EEG SignalsabstractEmotional information plays an important role in various multimedia applications. Movies, as a widely available form of multimedia content, can induce multiple positive emotions and stimulate people's pursuit of a better life. Different from negative emotions, positive emotions are highly correlated and difficult to distinguish in the emotional space. Since different positive emotions are often induced simultaneously by movies, traditional single-target or multi-class methods are not suitable for the classification of movie-induced positive emotions. In this paper, we proposeTransEEG, a model for multi-label positive emotion classification from a viewer's brain activities when watching emotional movies. The key features ofTransEEGinclude (1) explicitly modeling the spatial correlation and temporal dependencies of multi-channel EEG signals using the Transformer structure based model, which effectively addresses long-distance dependencies, (2) exploiting the label-label correlations to guide the discriminative EEG representation learning, for that we design an Inter-Emotion Mask for guiding the Multi-Head Attention to learn the inter-emotion correlations, and (3) constructing an attention score vector from the representation-label correlation matrix to refine emotion-relevant EEG features. To evaluate the ability of our model for multi-label positive emotion classification, we demonstrate our model on a state-of-the-art positive emotion database CPED. Extensive experimental results show that our proposed method achieves superior performance over the competitive approaches. Xiaobing Du, Xiaoming Deng 0001, Hangyu Qin, Yezhi Shu, Fang Liu 0035, Guozhen Zhao, Yukun Lai, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
IEEE Trans. Affect. Comput. | 9 |
| 2023 | Emotion Distribution Learning Based on Peripheral Physiological SignalsabstractEmotion analysis based on peripheral physiological signals has attracted increasing attention recently in affective computing. Previous works usually predict emotional states using a single emotion label for each discrete time. However, in real-world scenarios, it is not sufficient due to the fact that the real-world emotional state is usually a mixture of basic emotions. In this paper, we formulate the emotion analysis as an emotion distribution learning (EDL) problem and make two contributions. First, we establish a standardised dataset containing four negative emotions (anger, disgust, sadness, fear) and three positive emotions (tenderness, joy, amusement), which could be a useful benchmark for the EDL task. Second, we propose an emotion distribution prediction system which has the following distinct characteristics: (1) after processing raw peripheral physiological signals, we compute totally 89 representative features from four channels, i.e., GSR, SKT, ECG and HR, (2) an adaptive feature selection strategy based on recursive feature elimination (RFE) is used to select the most significant features in our EDL task, and (3) we design a dedicated EDL model based on convolution neural networks that takes information from both the feature correlation and the time domain into consideration. Experiments were conducted to validate our proposed system, and the results indicated that (1) the proposed feature selection strategy effectively selects significant features and improves algorithmic performance, and (2) the proposed EDL model can obtain good results in terms of six evaluation measures and outperform existing methods. Yezhi Shu, Niqi Liu, Guozhen Zhao, Yong-Jin Liu 0001 |
IEEE Trans. Affect. Comput. | 6 |
| 2023 | SparseDGCNN: Recognizing Emotion From Multichannel EEG SignalsabstractEmotion recognition from EEG signals has attracted much attention in affective computing. Recently, a novel dynamic graph convolutional neural network (DGCNN) model was proposed, which simultaneously optimized the network parameters and a weighted graph$G$characterizing the strength of functional relation between each pair of two electrodes in the EEG recording equipment. In this article, we propose a sparse DGCNN model which modifies DGCNN by imposing a sparseness constraint on$G$and improves the emotion recognition performance. Our work is based on an important observation: the tomography study reveals that different brain regions sampled by EEG electrodes may be related to different functions of the brain and then the functional relations among electrodes are possibly highly localized and sparse. However, introducing sparseness constraint into the graph$G$makes the loss function of sparse DGCNN non-differentiable at some singular points. To ensure that the training process of sparse DGCNN converges, we apply the forward-backward splitting method. To evaluate the performance of sparse DGCNN, we compare it with four representative recognition methods (SVM, DBN, GELM and DGCNN). In addition to comparing different recognition methods, our experiments also compare different features and spectral bands, including EEG features in time-frequency domain (DE, PSD, DASM, RASM, ASM and DCAU on different bands) extracted from four representative EEG datasets (SEED, DEAP, DREAMER, and CMEED). The results show that (1) sparse DGCNN has consistently better accuracy than representative methods and has a good scalability, and (2) DE, PSD, and ASM features on$\gamma$band convey most discriminative emotional information, and fusion of separate features and frequency bands can improve recognition performance. Minjing Yu, Yong-Jin Liu 0001, Guozhen Zhao, Dan Zhang 0014, Wenming Zheng |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | CPED: A Chinese Positive Emotion Database for Emotion Elicitation and AnalysisabstractPositive emotions are of great significance to people's daily life, such as human-computer/robot interaction. However, the structure of extensive positive emotions is not clear yet and effective standardized inducing materials containing as many positive emotional categories as possible are lacking. Thus, this article aims to establish a Chinese positive emotion database (CPED) to (1) effectively elicit positive emotion categories as many as possible, (2) provide both the subjective feelings of different positive emotions and a corresponding peripheral physiological database, and (3) explore the structure and framework of positive emotion categories. 42 video clips of 16 positive emotion categories were screened from 1000+ online clips. Then a total of 312 participants watched and rated these video clips during which GSR and PPG signals were recorded. 34 video clips that met hit rate and intensity standards were systemically clustered into four emotion categories (empathy, fun, creativity and esteem). Eventually, 22 film clips of these four major categories formed the CPED database. A total of 84 features from GSR and PPG signals were extracted and entered into RF, SVM, DBN and LSTM classifiers that serves as baseline classification methods. A classification accuracy of 44.66 percent for four major categories of positive emotions was achieved. Guozhen Zhao, Yezhi Shu, Yan Ge 0007, Dan Zhang 0014, Yong-Jin Liu 0001, Xianghong Sun |
IEEE Trans. Affect. Comput. | 6 |
| 2023 | Multi-Target Positive Emotion Recognition From EEG SignalsabstractCompared with the widely studied negative emotions in which different classes are easy to distinguish, nowadays less attention is paid to the recognition of positive emotions that are not fully independent. In this article, we propose to recognize multiple continuous positive emotions that exhibit statistical dependencies using multi-target regression — by analyzing brain activities when an individual watches emotional film clips — and explore the neural representation of different positive emotions. Thirty-seven participants volunteered to participate in our study, in which their brain activities were recorded when watching five selected film clips (corresponding to five positive emotions: amusement, happiness, romance, tenderness and warmth). First, 150 well-known power features extracted from Electroencephalography (EEG) signals and 105 multimedia content analysis features were collected as the pool of candidate features. Second, based on the collected features, we propose to use a linear model (linear regression) and a nonlinear model (long short-term memory network, LSTM) to predict the percentage of five positive emotions. Then, percentage values were converted to ranking numbers and Kendall rank correlation coefficients were calculated. Our results showed that (1) ensemble of regressor chains (ERC) using LSTM as unit regressor obtained both the best regression results (with lowest RMSE = 8.325 and highest$\text{R }^{2} = 0.346$) and the best Kendall rank correlation coefficient (0.165) on EEG features merely, and (2) selective features from alpha frequency bands of EEG signals could represent different positive emotions. These results demonstrate the effectiveness of selective EEG features on recognizing different positive emotions. Guozhen Zhao, Dan Zhang 0014, Yong-Jin Liu 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2023 | Fine-Grained Video Retrieval With Scene SketchesabstractBenefiting from the intuitiveness and naturalness of sketch interaction, sketch-based video retrieval (SBVR) has received considerable attention in the video retrieval research area. However, most existing SBVR research still lacks the capability of accurate video retrieval with fine-grained scene content. To address this problem, in this paper we investigate a new task, which focuses on retrieving the target video by utilizing a fine-grained storyboard sketch depicting the scene layout and major foreground instances' visual characteristics (e.g., appearance, size, pose, etc.) of video; we call such a task "fine-grained scene-level SBVR". The most challenging issue in this task is how to perform scene-level cross-modal alignment between sketch and video. Our solution consists of two parts. First, we construct a scene-level sketch-video dataset called SketchVideo, in which sketch-video pairs are provided and each pair contains a clip-level storyboard sketch and several keyframe sketches (corresponding to video frames). Second, we propose a novel deep learning architecture called Sketch Query Graph Convolutional Network (SQ-GCN). In SQ-GCN, we first adaptively sample the video frames to improve video encoding efficiency, and then construct appearance and category graphs to jointly model visual and semantic alignment between sketch and video. Experiments show that our fine-grained scene-level SBVR framework with SQ-GCN architecture outperforms the state-of-the-art fine-grained retrieval methods. The SketchVideo dataset and SQ-GCN code are available in the project webpage https://iscas-mmsketch.github.io/FG-SL-SBVR/. Ran Zuo, Xiaoming Deng 0001, Yukun Lai, Fang Liu 0035, CuiXia Ma, Hao Wang 0005, Yong-Jin Liu 0001, Hongan Wang |
IEEE Trans. Image Process. | 9 |
| 2023 | Positional Attention Guided Transformer-Like Architecture for Visual Question AnsweringabstractTransformer architectures have recently been introduced into the field of visual question answering (VQA), due to their powerful capabilities of information extraction and fusion. However, existing Transformer-like models, including models using a single Transformer structure and large-scale pre-training generic visual-linguistic models, do not fully utilize both positional information of words in questions and positional information of objects in images, which are shown in this paper to be crucial in VQA tasks. To address this challenge, we propose a novel positional attention guided Transformer-like architecture, which can adaptively extracts positional information within and across the visual and language modalities, and use this information to guide high-level interactions in inter- and intra-modality information flows. In particular, we design and assemble three positional attention modules into a single Transformer-like model MCAN. We show that the positional information introduced in intra-modality interaction can adaptively modulate inter-modality interaction according to different inputs, which plays an important role for visual reasoning. Experimental results demonstrate that our model outperforms the state-of-the-art models and is particularly good at handling object counting questions. Overall, our model achieves the accuracy of 70.10%, 71.27%, and 71.52% on the datasets of COCO-QA, VQA v1.0 test-std and VQA v2.0 test-std, respectively. Aihua Mao, Ken Lin, Jun Xuan, Yong-Jin Liu 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Audio-Driven Talking Face Video Generation With Dynamic Convolution KernelsabstractIn this paper, we present a dynamic convolution kernel (DCK) strategy for convolutional neural networks. Using a fully convolutional network with the proposed DCKs, high-quality talking-face video can be generated from multi-modal sources (i.e., unmatched audio and video) in real time, and our trained model is robust to different identities, head postures, and input audios. Our proposed DCKs are specially designed for audio-driven talking face video generation, leading to a simple yet effective end-to-end system. We also provide a theoretical analysis to interpret why DCKs work. Experimental results show that our method can generate high-quality talking-face video with background at 60 fps. Comparison and evaluation between our method and the state-of-the-art methods demonstrate the superiority of our method. Zipeng Ye, Mengfei Xia, Ran Yi 0002, Juyong Zhang, Yukun Lai, Xuwei Huang, Guo-Xin Zhang, Yong-Jin Liu 0001 |
IEEE Trans. Multim. | 8 |
| 2023 | Predicting Personalized Head Movement From Short Video and Speech SignalabstractAudio-driven talking face video generation has attracted much attention recently. However, few existing works pay attention to machine learning of talking head movement, especially based on the phonetic study. Observing that real-world talking faces often accompany natural head movement, in this paper, we model the relation between speech signal and talking head movement, which is a typical one-to-many mapping problem. To solve this problem, we propose a novel two-step mapping strategy: (1) in the first step, we train an encoder that predicts a head motion behavior pattern (modeled as a feature vector) from the head motion sequence of a short video of 10–15 seconds, and (2) in the second step, we train a decoder that predict a unique head motion sequence from both the motion behavior pattern and the auditory features of an arbitrary speech signal. Based on the proposed mapping strategy, we build a deep neural network model that takes a speech signal of a source person and a short video of a target person as input, and outputs a synthesized high-fidelity talking face video with personalized head pose. Extensive experiments and a user study show that our method can generate high-quality personalized head movement in synthesized talking face videos, and meanwhile, has comparable facial animation quality (e.g., lip synchronization and expression) with the state-of-the-art methods. Ran Yi 0002, Zipeng Ye, Zhiyao Sun, Juyong Zhang, Guo-Xin Zhang, Pengfei Wan 0001, Hujun Bao, Yong-Jin Liu 0001 |
IEEE Trans. Multim. | 8 |
| 2023 | PU-Flow: A Point Cloud Upsampling Network With Normalizing FlowsabstractPoint cloud upsampling aims to generate dense point clouds from given sparse ones, which is a challenging task due to the irregular and unordered nature of point sets. To address this issue, we present a novel deep learning-based model, called PU-Flow, which incorporates normalizing flows and weight prediction techniques to produce dense points uniformly distributed on the underlying surface. Specifically, we exploit the invertible characteristics of normalizing flows to transform points between euclidean and latent spaces and formulate the upsampling process as ensemble of neighbouring points in a latent space, where the ensemble weights are adaptively learned from local geometric context. Extensive experiments show that our method is competitive and, in most test cases, it outperforms state-of-the-art methods in terms of reconstruction quality, proximity-to-surface accuracy, and computation efficiency. The source code will be publicly available at https://github.com/unknownue/puflow. Aihua Mao, Zihui Du, Junhui Hou, Yaqi Duan, Yong-Jin Liu 0001, Ying He 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | STD-Net: Structure-Preserving and Topology-Adaptive Deformation Network for Single-View 3D Reconstructionabstract3D reconstruction from single-view images is a long-standing research problem. There have been various methods based on point clouds and volumetric representations. In spite of success in 3D models generation, it is quite challenging for these approaches to deal with models with complex topology and fine geometric details. Thanks to the recent advance of deep shape representations, learning the structure and detail representation using deep neural networks is a promising direction. In this article, we propose a novel approach named STD-Net to reconstruct 3D models utilizing mesh representation that is well suited for characterizing complex structures and geometry details. Our method consists of (1) an auto-encoder network for recovering the structure of an object with bounding box representation from a single-view image; (2) a topology-adaptive GCN for updating vertex position for meshes of complex topology; and (3) a unified mesh deformation block that deforms the structural boxes into structure-aware meshes. Evaluation on ShapeNet and PartNet shows that STD-Net has better performance than state-of-the-art methods in reconstructing complex structures and fine geometric details. Aihua Mao, Canglan Dai, Jie Yang 0038, Lin Gao 0004, Ying He 0001, Yong-Jin Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2023 | Yarn-Level Simulation of Hygroscopicity of Woven TextilesabstractSimulating liquid-textile interaction has received great attention in computer graphics recently. Most existing methods take textiles as particles or parameterized meshes. Although these methods can generate visually pleasing results, they cannot simulate water content at a microscopic level due to the lack of geometrically modeling of textile's anisotropic structure. In this paper, we develop a method for yarn-level simulation of hygroscopicity of textiles and evaluate it using various quantitative metrics. We model textiles in a fiber-yarn-fabric multi-scale manner and consider the dynamic coupled physical mechanisms of liquid spreading, including wetting, wicking, moisture sorption/desorption, and transient moisture-heat transfer in textiles. Our method can accurately simulate liquid spreading on textiles with different fiber materials and geometrical structures with consideration of air temperatures and humidity conditions. It visualizes the hygroscopicity of textiles to demonstrate their moisture management ability. We conduct qualitative and quantitative experiments to validate our method and explore various factors to analyze their influence on liquid spreading and hygroscopicity of textiles. Aihua Mao, Chaoqiang Xie, Huamin Wang 0001, Yong-Jin Liu 0001, Guiqing Li, Ying He 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | 3D-CariGAN: An End-to-End Solution to 3D Caricature Generation From Normal Face PhotosabstractCaricature is a type of artistic style of human faces that attracts considerable attention in the entertainment industry. So far a few 3D caricature generation methods exist and all of them require some caricature information (e.g., a caricature sketch or 2D caricature) as input. This kind of input, however, is difficult to provide by non-professional users. In this paper, we propose an end-to-end deep neural network model that generates high-quality 3D caricatures directly from a normal 2D face photo. The most challenging issue for our system is that the source domain of face photos (characterized by normal 2D faces) is significantly different from the target domain of 3D caricatures (characterized by 3D exaggerated face shapes and textures). To address this challenge, we: (1) build a large dataset of 5,343 3D caricature meshes and use it to establish a PCA model in the 3D caricature shape space; (2) reconstruct a normal full 3D head from the input face photo and use its PCA representation in the 3D caricature shape space to establish correspondences between the input photo and 3D caricature shape; and (3) propose a novel character loss and a novel caricature loss based on previous psychological studies on caricatures. Experiments including a novel two-level user study show that our system can generate high-quality 3D caricatures directly from normal face photos. Zipeng Ye, Mengfei Xia, Yanan Sun 0006, Ran Yi 0002, Minjing Yu, Juyong Zhang, Yukun Lai, Yong-Jin Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2023 | Stroke-based semantic segmentation for scene-level free-hand sketches
Xiaoming Deng 0001, Jinyao Li, Yukun Lai, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
Vis. Comput. | 6 |
| 2022 | PD-Flow: A Point Cloud Denoising Framework with Normalizing Flows
Aihua Mao, Zihui Du, Yu-Hui Wen, Jun Xuan, Yong-Jin Liu 0001 |
ECCV (3) | 5 |
| 2022 | Audio-Driven Stylized Gesture Generation with Flow-Based Model
Yu-Hui Wen, Yanan Sun 0006, Ying He 0001, Yaoyuan Wang, Weihua He, Yong-Jin Liu 0001 |
ECCV (5) | 8 |
| 2022 | ParticleSfM: Exploiting Dense Point Trajectories for Localizing Moving Cameras in the Wild
Wang Zhao 0001, Shaohui Liu, Hengkai Guo, Wenping Wang 0001, Yong-Jin Liu 0001 |
ECCV (32) | 5 |
| 2022 | A Double Branch Next-Best-View Network and Novel Robot System for Active Object ReconstructionabstractNext best view (NBV) is a technology that finds the best view sequence for sensor to perform scanning based on partial information, which is the core part for robot active reconstruction. Traditional works are mostly based on the evaluation of candidate views through time-consuming volu-metric transformation and ray casting, which heavily limits the applications of NBV. Recent deep learning based NBV methods aim to approximately learn the evaluation function by large-scale training, and improve both the effectiveness and efficiency of NBV. However, these methods force the network to regress the exact groundtruth value of each candidate view, which is much harder than simply ranking all the candidate views. Besides, most previous NBV works assume perfect sensing and perform in simulation environments, lacking real application abilities. In this paper, we propose a novel double branch NBV network, DB-NBV, to utilize the ranking process together with the evaluation process. We further design a real NBV robot and a pipeline to conduct real active reconstruction. Experiments on both simulation and real robot show that our method achieves the best performance and can be applied to real application with high accuracy and speed. Yiheng Han, Irvin Haozhe Zhan, Wang Zhao 0001, Yong-Jin Liu 0001 |
ICRA | 4 |
| 2022 | Generating Smooth and Facial-Details-Enhanced Talking Head Video: A Perspective of Pre and Post ProcessesabstractTalking head video generation has received increasing attention recently. So far the quality (especially the facial details) of the videos output from state-of-the-art deep learning methods is limited by either the quality of training data or the performance of generators, and needs to be further improved. In this paper, we propose a data pre- and post- processing strategy based on a key observation: generating talking head video from multi-modal input is a challenging problem and generating smooth video with fine facial details makes the problem even harder. Then we propose to decompose the problem solution into a main deep model, a pre- and a post- processing. The main deep model generates a reasonably good talking face video, with the aid of a pre-process, which also contributes to a post-process for restoring smooth and fine facial details in the final video. In particular, our main deep model reconstructs a 3D face from an input reference frame, and then uses an AudioNet to generate a sequence of facial expression coefficients with an input audio clip. To ensure final facial details in the generated video, we sample the original texture from the reference frame in the pre-process with the aid of reconstructed 3D face and a predefined UV map. Accordingly, in the post-process, we smooth the expression coefficients of adjacent frames to alleviate jitters and apply a pretrained face restoration module to recover the fine facial details. Experimental results and ablation study show the advantage of our proposed method. Tian Lv, Yu-Hui Wen, Zhiyao Sun, Zipeng Ye, Yong-Jin Liu 0001 |
ACM Multimedia | 5 |
| 2022 | A Mixture Of Surprises for Unsupervised Reinforcement LearningabstractUnsupervised reinforcement learning aims at learning a generalist policy in a reward-free manner for fast adaptation to downstream tasks. Most of the existing methods propose to provide an intrinsic reward based on surprise. Maximizing or minimizing surprise drives the agent to either explore or gain control over its environment. However, both strategies rely on a strong assumption: the entropy of the environment's dynamics is either high or low. This assumption may not always hold in real-world scenarios, where the entropy of the environment's dynamics may be unknown. Hence, choosing between the two objectives is a dilemma. We propose a novel yet simple mixture of policies to address this concern, allowing us to optimize an objective that simultaneously maximizes and minimizes the surprise. Concretely, we train one mixture component whose objective is to maximize the surprise and another whose objective is to minimize the surprise. Hence, our method does not make assumptions about the entropy of the environment's dynamics. We call our method a $\textbf{M}\text{ixture }\textbf{O}\text{f }\textbf{S}\text{urprise}\textbf{S}$ (MOSS) for unsupervised reinforcement learning. Experimental results show that our simple method achieves state-of-the-art performance on the URLB benchmark, outperforming previous pure surprise maximization-based objectives. Our code is available at: https://github.com/LeapLabTHU/MOSS. Andrew Zhao, Matthieu Lin, Yangguang Li 0001, Yong-Jin Liu 0001, Gao Huang 0001 |
NeurIPS | 4 |
| 2022 | NPRportrait 1.0: A three-level benchmark for non-photorealistic rendering of portraitsabstractRecently, there has been an upsurge of activity in image-based non-photorealistic rendering (NPR), and in particular portrait image stylisation, due to the advent of neural style transfer (NST). However, the state of performance evaluation in this field is poor, especially compared to the norms in the computer vision and machine learning communities. Unfortunately, the task of evaluating image stylisation is thus far not well defined, since it involves subjective, perceptual, and aesthetic aspects. To make progress towards a solution, this paper proposes a new structured, three-level, benchmark dataset for the evaluation of stylised portrait images. Rigorous criteria were used for its construction, and its consistency was validated by user studies. Moreover, a new methodology has been developed for evaluating portrait stylisation algorithms, which makes use of the different benchmark levels as well as annotations provided by user studies regarding the characteristics of the faces. We perform evaluation for a wide variety of image stylisation methods (both portrait-specific and general purpose, and also both traditional NPR approaches and NST) using the new benchmark dataset. Paul L. Rosin, Yukun Lai, David Mould, Ran Yi 0002, Itamar Berger, Lars Doyle, Seungyong Lee 0001, Chuan Li 0001, Yong-Jin Liu 0001, Amir Semmo, Ariel Shamir, Minjung Son 0001, Holger Winnemöller |
Comput. Vis. Media | 9 |
| 2022 | Video-Based Facial Micro-Expression Analysis: A Survey of Datasets, Features and AlgorithmsabstractUnlike the conventional facial expressions, micro-expressions are involuntary and transient facial expressions capable of revealing the genuine emotions that people attempt to hide. Therefore, they can provide important information in a broad range of applications such as lie detection, criminal detection, etc. Since micro-expressions are transient and of low intensity, however, their detection and recognition is difficult and relies heavily on expert experiences. Due to its intrinsic particularity and complexity, video-based micro-expression analysis is attractive but challenging, and has recently become an active area of research. Although there have been numerous developments in this area, thus far there has been no comprehensive survey that provides researchers with a systematic overview of these developments with a unified evaluation. Accordingly, in this survey paper, we first highlight the key differences between macro- and micro-expressions, then use these differences to guide our research survey of video-based micro-expression analysis in a cascaded structure, encompassing the neuropsychological basis, datasets, features, spotting algorithms, recognition algorithms, applications and evaluation of state-of-the-art approaches. For each aspect, the basic techniques, advanced developments and major challenges are addressed and discussed. Furthermore, after considering the limitations of existing micro-expression datasets, we present and release a new dataset — calledmicro-and-macro expression warehouse(MMEW) — containing more video samples and more labeled emotion types. We then perform a unified comparison of representative methods on CAS(ME)$^2$for spotting, and on MMEW and SAMM for recognition, respectively. Finally, some potential future research directions are explored and outlined. Xianye Ben, Junping Zhang, Kidiyo Kpalma, Weixiao Meng 0001, Yong-Jin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | An Efficient LSTM Network for Emotion Recognition From Multichannel EEG SignalsabstractMost previous EEG-based emotion recognition methods studied hand-crafted EEG features extracted from different electrodes. In this article, we study the relation among different EEG electrodes and propose a deep learning method to automatically extract the spatial features that characterize the functional relation between EEG signals at different electrodes. Our proposed deep model is calledATtention-basedLSTMwithDomainDiscriminator (ATDD-LSTM), a model based on Long Short-Term Memory (LSTM) for emotion recognition that can characterize nonlinear relations among EEG signals of different electrodes. To achieve state-of-the-art emotion recognition performance, the architecture of ATDD-LSTM has two distinguishing characteristics: (1) By applying the attention mechanism to the feature vectors produced by LSTM, ATDD-LSTM automatically selects suitable EEG channels for emotion recognition, which makes the learned model concentrate on the emotion related channels in response to a given emotion; (2) To minimize the significant feature distribution shift between different sessions and/or subjects, ATDD-LSTM uses a domain discriminator to modify the data representation space and generate domain-invariant features. We evaluate the proposed ATDD-LSTM model on three public EEG emotional databases (DEAP, SEED and CMEED) for emotion recognition. The experimental results demonstrate that our ATDD-LSTM model achieves superior performance on subject-dependent (for the same subject), subject-independent (for different subjects) and cross-session (for the same subject) evaluation. Xiaobing Du, CuiXia Ma, Jinyao Li, Yukun Lai, Guozhen Zhao, Xiaoming Deng 0001, Yong-Jin Liu 0001, Hongan Wang |
IEEE Trans. Affect. Comput. | 8 |
| 2022 | PPR-Net++: Accurate 6-D Pose Estimation in Stacked ScenariosabstractMost supervised learning-based pose estimation methods for stacked scenes are trained on massive synthetic datasets. In most cases, the challenge is that the learned network on the training dataset is no longer optimal on the testing dataset. To address this problem, we propose a pose regression network PPR-Net++. It transforms each scene point into a point in the centroid space, followed by a clustering process and a voting process. In the training phase, a mapping function between the network’s critical parameter (i.e., the bandwidth of the clustering algorithm) and the compactness of the centroid distributions is obtained. This function is used to adapt the bandwidth between centroid distributions of two different domains. In addition, to further improve the pose estimation accuracy, the network also predicts the confidence of each point, based on its visibility and pose error. Only the points with high confidence have the right to vote for the final object pose. In experiments, our method is trained on the IPA synthetic dataset and compared with the state-of-the-art algorithm. When tested with the public synthetic Siléane dataset, our method is better in all eight objects, where five of them are improved by more than 5% in average precision (AP). On IPA real dataset, our method outperforms a large margin by 20%. This lays a solid foundation for robot grasping in industrial scenarios. Note to Practitioners—Our work is motivated by industrial product assembly based on robot grasping. The industrial parts are usually manufactured by numerical machines and piled in bins. Our method can estimate the poses of visible parts accurately. A pose of a part includes its centroid and spatial orientations. Combined with a depth camera, this algorithm allows an industrial robot to understand complex stacked scenes. We improve the pose estimation accuracy in order to assemble parts with robot grasping, without an additional pose adjuster. Our network can learn from a synthetic dataset and apply it to real-world data, without a significant accuracy drop. The synthetic dataset can be obtained easily by computer simulation programs, so the training data are sufficient. Experiments demonstrate that our method outperforms the state-of-the-art pose estimation approaches. Long Zeng 0001, Wei Jie Lv, Zhi-Kai Dong, Yong-Jin Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2022 | SketchMaker: Sketch Extraction and Reuse for Interactive Scene Sketch CompositionabstractSketching is an intuitive and simple way to depict sciences with various object form and appearance characteristics. In the past few years, widely available touchscreen devices have increasingly made sketch-based human-AI co-creation applications popular. One key issue of sketch-oriented interaction is to prepare input sketches efficiently by non-professionals because it is usually difficult and time-consuming to draw an ideal sketch with appropriate outlines and rich details, especially for novice users with no sketching skills. Thus, sketching brings great obstacles for sketch applications in daily life. On the other hand, hand-drawn sketches are scarce and hard to collect. Given the fact that there are several large-scale sketch datasets providing sketch data resources, but they usually have a limited number of objects and categories in sketch, and do not support users to collect new sketch materials according to their personal preferences. In addition, few sketch-related applications support the reuse of existing sketch elements. Thus, knowing how to extract sketches from existing drawings and effectively re-use them in interactive scene sketch composition will provide an elegant way for sketch-based image retrieval (SBIR) applications, which are widely used in various touch screen devices. In this study, we first conduct a study on current SBIR to better understand the main requirements and challenges in sketch-oriented applications. Then we develop the SketchMaker as an interactive sketch extraction and composition system to help users generate scene sketches via reusing object sketches in existing scene sketches with minimal manual intervention. Moreover, we demonstrate how SBIR improves from composited scene sketches to verify the performance of our interactive sketch processing system. We also include a sketch-based video localization task as an alternative application of our sketch composition scheme. Our pilot study shows that our system is effective and efficient, and provides a way to promote practical applications of sketches. Fang Liu 0035, Xiaoming Deng 0001, Jian-Cheng Song, Yukun Lai, Yong-Jin Liu 0001, Hao Wang 0005, CuiXia Ma, Sheng Feng Qin, Hongan Wang |
ACM Trans. Interact. Intell. Syst. | 5 |
| 2022 | SceneSketcher-v2: Fine-Grained Scene-Level Sketch-Based Image Retrieval Using Adaptive GCNsabstractSketch-based image retrieval (SBIR) is a long-standing research topic in computer vision. Existing methods mainly focus on category-level or instance-level image retrieval. This paper investigates the fine-grained scene-level SBIR problem where a free-hand sketch depicting a scene is used to retrieve desired images. This problem is useful yet challenging mainly because of two entangled facts: 1) achieving an effective representation of the input query data and scene-level images is difficult as it requires to model the information across multiple modalities such as object layout, relative size and visual appearances, and 2) there is a great domain gap between the query sketch input and target images. We present SceneSketcher-v2, a Graph Convolutional Network (GCN) based architecture to address these challenges. SceneSketcher-v2 employs a carefully designed graph convolution network to fuse the multi-modality information in the query sketch and target images and uses a triplet training process and end-to-end training manner to alleviate the domain gap. Extensive experiments demonstrate SceneSketcher-v2 outperforms state-of-the-art scene-level SBIR models with a significant margin. Fang Liu 0035, Xiaoming Deng 0001, Changqing Zou, Yukun Lai, Ran Zuo, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
IEEE Trans. Image Process. | 8 |
| 2022 | E-ffective: A Visual Analytic System for Exploring the Emotion and Effectiveness of Inspirational SpeechesabstractWhat makes speeches effective has long been a subject for debate, and until today there is broad controversy among public speaking experts about what factors make a speech effective as well as the roles of these factors in speeches. Moreover, there is a lack of quantitative analysis methods to help understand effective speaking strategies. In this paper, we propose E-ffective, a visual analytic system allowing speaking experts and novices to analyze both the role of speech factors and their contribution in effective speeches. From interviews with domain experts and investigating existing literature, we identified important factors to consider in inspirational speeches. We obtained the generated factors from multi-modal data that were then related to effectiveness data. Our system supports rapid understanding of critical factors in inspirational speeches, including the influence of emotions by means of novel visualization methods and interaction. Two novel visualizations include E-spiral (that shows the emotional shifts in speeches in a visually compact way) and E-script (that connects speech content with key speech delivery information). In our evaluation we studied the influence of our system on experts' domain knowledge about speech factors. We further studied the usability of the system by speaking novices and experts on assisting analysis of inspirational speech effectiveness. Kevin T. Maher, Jian-Cheng Song, Xiaoming Deng 0001, Yukun Lai, CuiXia Ma, Hao Wang 0005, Yong-Jin Liu 0001, Hongan Wang |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2022 | GAN-Based Multi-Style Photo CartoonizationabstractCartoon is a common form of art in our daily life and automatic generation of cartoon images from photos is highly desirable. However, state-of-the-art single-style methods can only generate one style of cartoon images from photos and existing multi-style image style transfer methods still struggle to produce high-quality cartoon images due to their highly simplified and abstract nature. In this article, we propose a novel multi-style generative adversarial network (GAN) architecture, called MS-CartoonGAN, which can transform photos into multiple cartoon styles. MS-CartoonGAN uses only unpaired photos and cartoon images of multiple styles for training. To achieve this, we propose to use (1) a hierarchical semantic loss with sparse regularization to retain semantic content and recover flat shading in different abstract levels, (2) a new edge-promoting adversarial loss for producing fine edges, and (3) a style loss to enhance the difference between output cartoon styles and make training process more stable. We also develop a multi-domain architecture, where the generator consists of a shared encoder and multiple decoders for different cartoon styles, along with multiple discriminators for individual styles. By observing that cartoon images drawn by different artists have their unique styles while sharing some common characteristics, our shared network architecture exploits the common characteristics of cartoon styles, achieving better cartoonization and being more efficient than single-style cartoonization. We show that our multi-domain architecture can theoretically guarantee to output desired multiple cartoon styles. Through extensive experiments including a user study, we demonstrate the superiority of the proposed method, outperforming state-of-the-art single-style and multi-style image style transfer methods. Yezhi Shu, Ran Yi 0002, Mengfei Xia, Zipeng Ye, Wang Zhao 0001, Yukun Lai, Yong-Jin Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2021 | Autoregressive Stylized Motion Synthesis With Generative FlowabstractMotion style transfer is an important problem in many computer graphics and computer vision applications, including human animation, games, and robotics. Most existing deep learning methods for this problem are supervised and trained by registered motion pairs. In addition, these methods are often limited to yielding a deterministic output, given a pair of style and content motions. In this paper, we propose an unsupervised approach for motion style transfer by synthesizing stylized motions autoregressively using a generative flow model $\mathcal{M}$. $\mathcal{M}$ is trained to maximize the exact likelihood of a collection of unlabeled motions, based on an autoregressive context of poses in previous frames and a control signal representing the movement of a root joint. Thanks to invertible flow transformations, latent codes that encode deep properties of motion styles are efficiently inferred by $\mathcal{M}$. By combining the latent codes (from an input style motion S) with the autoregressive context and control signal (from an input content motion C), $\mathcal{M}$ outputs a stylized motion which transfers style from S to C. Moreover, our model is probabilistic and is able to generate various plausible motions with a specific style. We evaluate the proposed model on motion capture datasets containing different human motion styles. Experiment results show that our model outperforms the state-of-the-art methods, despite not requiring manually labeled training data. Yu-Hui Wen, Hongbo Fu 0001, Lin Gao 0004, Yanan Sun 0006, Yong-Jin Liu 0001 |
CVPR | 6 |
| 2021 | AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisabstractGenerating high-fidelity talking head video by fitting with the input audio sequence is a challenging problem that receives considerable attentions recently. In this paper, we address this problem with the aid of neural scene representation networks. Our method is completely different from existing methods that rely on intermediate representations like 2D landmarks or 3D face models to bridge the gap between audio input and video output. Specifically, the feature of input audio signal is directly fed into a conditional implicit function to generate a dynamic neural radiance field, from which a high-fidelity talking-head video corresponding to the audio signal is synthesized using volume rendering. Another advantage of our framework is that not only the head (with hair) region is synthesized as previous methods did, but also the upper body is generated via two individual neural radiance fields. Experimental results demonstrate that our novel framework can (1) produce high-fidelity and natural results, and (2) support free adjustment of audio signals, viewing directions, and background images. Code is available at https://github.com/YudongGuo/AD-NeRF. Sen Liang, Yong-Jin Liu 0001, Hujun Bao, Juyong Zhang |
ICCV | 4 |
| 2021 | A Confidence-based Iterative Solver of Depths and Surface Normals for Deep Multi-view StereoabstractIn this paper, we introduce a deep multi-view stereo (MVS) system that jointly predicts depths, surface normals and per-view confidence maps. The key to our approach is a novel solver that iteratively solves for per-view depth map and normal map by optimizing an energy potential based on the locally planar assumption. Specifically, the algorithm updates depth map by propagating from neigh-boring pixels with slanted planes, and updates normal map with local probabilistic plane fitting. Both two steps are monitored by a customized confidence map. This solver is not only effective as a post-processing tool for plane-based depth refinement and completion, but also differentiable such that it can be efficiently integrated into deep learning pipelines. Our multi-view stereo system employs multiple optimization steps of the solver over the initial prediction of depths and surface normals. The whole system can be trained end-to-end, decoupling the challenging problem of matching pixels within poorly textured regions from the cost-volume based neural network. Experimental results on ScanNet and RGB-D Scenes V2 demonstrate state-of-the-art performance of the proposed deep MVS system on multi-view depth estimation, with our proposed solver consistently improving the depth quality over both conventional and deep learning based MVS pipelines. Code is available at https://github.com/thuzhaowang/idn-solver. Wang Zhao 0001, Shaohui Liu, Yi Wei 0003, Hengkai Guo, Yong-Jin Liu 0001 |
ICCV | 5 |
| 2021 | Efficient SE(3) Reachability Map Generation via Interplanar Integration of Intra-planar ConvolutionsabstractConvolution has been used for fast computation of reachability maps, but it has high computational costs when performing SE(3) convolution operations for general joint arrangements in industrial robots and 3D workspace. Its application is also limited to planar robots, 2D workspace, or robots with special spatial arrangements for joints. In this paper, we find that the SE(3) convolution can be decomposed into a set of SE(2) convolutions, which significantly reduces the computational complexity when computing the reachability map of high-DOF robotic manipulators in the 3D workspace. We also leverage GPU parallel computing and Fast Fourier transform to further accelerate the computation procedure. We demonstrate the time efficiency and quality of our approach using a set of numerical experiments for constructing reachability maps and also present a multi-robot plant phenotyping system that uses the computed reachability map for efficient viewpoint selection and path planning. Yiheng Han, Jia Pan 0001, Mengfei Xia, Long Zeng 0001, Yong-Jin Liu 0001 |
ICRA | 5 |
| 2021 | ParametricNet: 6DoF Pose Estimation Network for Parametric Shapes in Stacked ScenariosabstractMost industrial parts are parametric and their special properties are not fully explored yet. This paper proposes a new 6DoF pose estimation network for parametric shapes in stacked scenarios (ParametricNet). It treats a parametric shape, instead of a part object, as a category. The keypoints of individual instances are learned with point- wise regression and Hough voting scheme, from which specific parameter values are calculated. Then, the template keypoints are obtained based on the computed parameter values and the parametric shape templates. Finally, the 6DoF pose is estimated by least-square fitting between the individual instance’s and the template’s keypoints & centroid. On the public Siléane dataset, the average of APs of ParametricNet is 96%, compared with 82% for the state-of-the-art method. In addition, a new parametric dataset with four shape templates is constructed, in which the evaluated learning and generalization abilities of ParametricNet outperform the state-of-the-art methods. In particular, for the less symmetric shape, the mAP is improved by over 20%, which is an obvious improvement. Real-world experiments show that our method can grasp parametric shapes with unknown parameter values in stacked scenarios. Long Zeng 0001, Wei Jie Lv, Yong-Jin Liu 0001 |
ICRA | 4 |
| 2021 | Line Drawings for Face Portraits From Photos Using Global and Local Structure Based GANsabstractDespite significant effort and notable success of neural style transfer, it remains challenging for highly abstract styles, in particular line drawings. In this paper, we propose APDrawingGAN++, a generative adversarial network (GAN) for transforming face photos to artistic portrait drawings (APDrawings), which addresses substantial challenges including highly abstract style, different drawing techniques for different facial features, and high perceptual sensitivity to artifacts. To address these, we propose a composite GAN architecture that consists of local networks (to learn effective representations for specific facial features) and a global network (to capture the overall content). We provide a theoretical explanation for the necessity of this composite GAN structure by proving that any GAN with a single generator cannot generate artistic styles like APDrawings. We further introduce a classification-and-synthesis approach for lips and hair where different drawing styles are used by artists, which applies suitable styles for a given input. To capture the highly abstract art form inherent in APDrawings, we address two challenging operations-(1) coping with lines with small misalignments while penalizing large discrepancy and (2) generating more continuous lines-by introducing two novel loss terms: one is a novel distance transform loss with nonlinear mapping and the other is a novel line continuity loss, both of which improve the line quality. We also develop dedicated data augmentation and pre-training to further improve results. Extensive experiments, including a user study, show that our method outperforms state-of-the-art methods, both qualitatively and quantitatively. Ran Yi 0002, Mengfei Xia, Yong-Jin Liu 0001, Yukun Lai, Paul L. Rosin |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Feature-Aware Uniform Tessellations on Video Manifold for Content-Sensitive SupervoxelsabstractOver-segmenting a video into supervoxels has strong potential to reduce the complexity of downstream computer vision applications. Content-sensitive supervoxels (CSSs) are typically smaller in content-dense regions (i.e., with high variation of appearance and/or motion) and larger in content-sparse regions. In this paper, we propose to compute feature-aware CSSs (FCSSs) that are regularly shaped 3D primitive volumes well aligned with local object/region/motion boundaries in video. To compute FCSSs, we map a video to a 3D manifold embedded in a combined color and spatiotemporal space, in which the volume elements of video manifold give a good measure of the video content density. Then any uniform tessellation on video manifold can induce CSS in the video. Our idea is that among all possible uniform tessellations on the video manifold, FCSS finds one whose cell boundaries well align with local video boundaries. To achieve this goal, we propose a novel restricted centroidal Voronoi tessellation method that simultaneously minimizes the tessellation energy (leading to uniform cells in the tessellation) and maximizes the average boundary distance (leading to good local feature alignment). Theoretically our method has an optimal competitive ratio O(1), and its time and space complexities are O(NK) and O(N+K) for computing K supervoxels in an N-voxel video. We also present a simple extension of FCSS to streaming FCSS for processing long videos that cannot be loaded into main memory at once. We evaluate FCSS, streaming FCSS and ten representative supervoxel methods on four video datasets and two novel video applications. The results show that our method simultaneously achieves state-of-the-art performance with respect to various evaluation criteria. Ran Yi 0002, Zipeng Ye, Wang Zhao 0001, Minjing Yu, Yukun Lai, Yong-Jin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | Inter-Brain EEG Feature Extraction and Analysis for Continuous Implicit Emotion Tagging During Video WatchingabstractHow to efficiently tag the emotional experience of multimedia contents is an important and challenging problem in the field of affective computing. This paper presents an EEG-based real-time emotion tagging approach, by extracting inter-brain features from a group of participants when they watch the same emotional video clips. First, the continuous subjective reports on both the arousal and valence dimensions of emotion were obtained by employing a three-round behavioral rating paradigm. Second, the inter-brain features were systematically explored in both spectral and temporal domain. Finally, regression analyses were performed to evaluate the effectiveness of inter-brain amplitude and phase features. The inter-brain amplitude feature showed significantly better prediction performance than the inter-brain phase feature, as well as another two conventional features (spectral power and inter-subject correlation). By combining the four types of features, regression values (R2) were obtained for the prediction of arousal (0.61 + 0.01) and valence (0.70 + 0.01), corresponding to prediction errors of 1.01 + 0.02 and 0.78 + 0.02 (unit on 9-point scales), respectively. The contributions of different electrodes and frequency bands were also analyzed. Our results show promising potentials of inter-brain EEG features in real-time emotion tagging applications. Yue Ding 0005, Zhenyi Xia, Yong-Jin Liu 0001, Dan Zhang 0014 |
IEEE Trans. Affect. Comput. | 4 |
| 2021 | Sparse MDMO: Learning a Discriminative Feature for Micro-Expression RecognitionabstractMicro-expressions are the rapid movements of facial muscles that can be used to reveal concealed emotions. Recognizing them from video clips has a wide range of applications and receives increasing attention recently. Among existing methods, the main directional mean optical-flow (MDMO) feature achieves state-of-the-art performance for recognizing spontaneous micro-expressions. For a video clip, the MDMO feature is computed by averaging a set of atomic features frame-by-frame. Despite its simplicity, the average operation in MDMO can easily lose the underlying manifold structure inherent in the feature space. In this paper we propose a sparse MDMO feature that learns an effective dictionary from a micro-expression video dataset. In particular, a new distance metric is proposed based on the sparsity of sample points in the MDMO feature space, which can efficiently reveal the underlying manifold structure. The proposed sparse MDMO feature is obtained by incorporating this new metric into the classic graph regularized sparse coding (GraphSC) scheme. We evaluate sparse MDMO and four representative features (LBP-TOP, STCLQP, MDMO and FDM) on three spontaneous micro-expression datasets (SMIC, CASME and CASME II). The results show that sparse MDMO outperforms these representative features. Yong-Jin Liu 0001, Bing-Jun Li, Yukun Lai |
IEEE Trans. Affect. Comput. | 1 |
| 2021 | GPU-Based Supervoxel Generation With a Novel Anisotropic MetricabstractVideo over-segmentation into supervoxels is an important pre-processing technique for many computer vision tasks. Videos are an order of magnitude larger than images. Most existing methods for generating supervovels are either memory- or time-inefficient, which limits their application in subsequent video processing tasks. In this paper, we present an anisotropic supervoxel method, which is memory-efficient and can be executed on the graphics processing unit (GPU). Therefore, our algorithm achieves good balance among segmentation quality, memory usage and processing time. In order to provide accurate segmentation for moving objects in video, we use the optical flow information to design a brand new non-Euclidean metric to calculate the anisotropic distances between seeds and voxels. To efficiently compute the anisotropic metric, we adjust the classic jump flooding algorithm (which is designed for parallel execution on the GPU) to generate anisotropic Voronoi tessellation in the combined color and spatio-temporal space. We evaluate our method and the representative supervoxel algorithms for their capability on segmentation performance, computation speed and memory efficiency. We also apply supervoxel results to the application of foreground propagation in videos to test the performance on solving practical problems. Experiments show that our algorithm is much faster than the existing methods, and achieves good balance on segmentation quality and efficiency. Zhonggui Chen, Yong-Jin Liu 0001, Junfeng Yao, Xiaohu Guo |
IEEE Trans. Image Process. | 3 |
| 2021 | Automatic Sitting Pose Generation for Ergonomic Ratings of ChairsabstractHuman poses play a critical role in human-centric product design. Despite considerable researches on pose synthesis and pose-driven product design, most of them adopt the simple stick figure model that captures only skeletons rather than real body geometries and do not link human poses to the environment (e.g., chairs for sitting). This paper focuses on user-tailored ergonomic design and rating of chairs using scanned human geometries. Fully utilizing the anthropometric information of the human models, our method considers more ergonomic guidelines of chair design (such as pressure distribution and support intensity) and links the geometry of 3D chair models and human-to-chair interactions into the pose deformation constraints of the human avatars. The core of our method is a pose generation algorithm which rigs the user's successive poses through coarse- and fine-level pose deformations. We define a non-linear energy function with contact, collision, and joint limit terms, and solve it using a hill-climbing algorithm. The fitting results allow us to quantitatively evaluate the chair model in terms of various ergonomic criteria. Our method is flexible and effective and can be applied to users with varying body shapes and a wide range of chairs. Moreover, the proposed technique can be easily extended to other furniture, such as desk, bed, and cabinet. Extensive evaluations and a user study demonstrate the efficiency and advantages of the proposed virtual fitting method. Given that our method avoids tedious on-site trying, facilitates the exploration/evaluation of various chair products, and provides valuable feedback for the designers and manufacturers to deliver customized products, it is ideal for online shopping of chairs. Aihua Mao, Zhenfeng Xie, Minjing Yu, Yong-Jin Liu 0001, Ying He 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2020 | Unpaired Portrait Drawing Generation via Asymmetric Cycle MappingabstractPortrait drawing is a common form of art with high abstraction and expressiveness. Due to its unique characteristics, existing methods achieve decent results only with paired training data, which is costly and time-consuming to obtain.In this paper, we address the problem of automatic transfer from face photos to portrait drawings with unpaired training data. We observe that due to the significant imbalance of information richness between photos and drawings, existing unpaired transfer methods such as CycleGAN tends to embed invisible reconstruction information indiscriminately in the whole drawings, leading to important facial features partially missing in drawings. To address this problem, we propose a novel asymmetric cycle mapping that enforces the reconstruction information to be visible (by a truncation loss) and only embedded in selective facial regions (by a relaxed forward cycle-consistency loss). Along with localized discriminators for the eyes, nose and lips, our method well preserves all important facial features in the generated portrait drawings. By introducing a style classifier and taking the style vector into account, our method can learn to generate portrait drawings in multiple styles using a single network. Extensive experiments show that our model outperforms state-of-the-art methods. Ran Yi 0002, Yong-Jin Liu 0001, Yukun Lai, Paul L. Rosin |
CVPR | 2 |
| 2020 | Towards Better Generalization: Joint Depth-Pose Learning Without PoseNetabstractIn this work, we tackle the essential problem of scale inconsistency for self supervised joint depth-pose learning. Most existing methods assume that a consistent scale of depth and pose can be learned across all input samples, which makes the learning problem harder, resulting in degraded performance and limited generalization in indoor environments and long-sequence visual odometry application. To address this issue, we propose a novel system that explicitly disentangles scale from the network estimation. Instead of relying on PoseNet architecture, our method recovers relative pose by directly solving fundamental matrix from dense optical flow correspondence and makes use of a two-view triangulation module to recover an up-to-scale 3D structure. Then, we align the scale of the depth prediction with the triangulated point cloud and use the transformed depth map for depth error computation and dense reprojection check. Our whole system can be jointly trained end-to-end. Extensive experiments show that our system not only reaches state-of-the-art performance on KITTI depth and flow estimation, but also significantly improves the generalization ability of existing self-supervised depth-pose learning methods under a variety of challenging scenarios, and achieves state-of-the-art results among self-supervised learning-based methods on KITTI Odometry and NYUv2 dataset. Furthermore, we present some interesting findings on the limitation of PoseNet-based relative pose estimation methods in terms of generalization ability. Code is available at https://github.com/B1ueber2y/TrianFlow. Wang Zhao 0001, Shaohui Liu, Yezhi Shu, Yong-Jin Liu 0001 |
CVPR | 4 |
| 2020 | SceneSketcher: Fine-Grained Image Retrieval with Scene Sketches
Fang Liu 0035, Changqing Zou, Xiaoming Deng 0001, Ran Zuo, Yukun Lai, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
ECCV (19) | 7 |
| 2020 | Configuration Space Decomposition for Learning-based Collision Checking in High-DOF RobotsabstractMotion planning for robots of high degrees-of-freedom (DOFs) is an important problem in robotics with sampling-based methods in configuration space $\mathcal{C}$ as one popular solution. Recently, machine learning methods have been introduced into sampling-based motion planning methods, which train a classifier to distinguish collision free subspace from in-collision subspace in $\mathcal{C}$. In this paper, we propose a novel configuration space decomposition method and show two nice properties resulted from this decomposition. Using these two properties, we build a composite classifier that works compatibly with previous machine learning methods by using them as the elementary classifiers. Experimental results are presented, showing that our composite classifier outperforms state-of-the-art single-classifier methods by a large margin. A real application of motion planning in a multi-robot system in plant phenotyping using three UR5 robotic arms is also presented. Yiheng Han, Wang Zhao 0001, Jia Pan 0001, Yong-Jin Liu 0001 |
IROS | 4 |
| 2020 | PC-NBV: A Point Cloud Based Deep Network for Efficient Next Best View PlanningabstractThe Next Best View (NBV) problem is important in the active robotic reconstruction. It enables the robot system to perform scanning actions in a reasonable view sequence, and fulfil the reconstruction task in an effective way. Previous works mainly follow the volumetric methods, which convert the point cloud information collected by sensors into a voxel representation space and evaluate candidate views through ray casting simulations to pick the NBV. However, the process of volumetric data transformation and ray casting is often time-consuming. To address this issue, in this paper, we propose a point cloud based deep neural network called PC-NBV to achieve efficient view planning without these computationally expensive operations. The PC-NBV network takes the raw point cloud data and current view selection states as input, and then directly predicts the information gain of all candidate views. By avoiding costly data transformation and ray casting, and utilizing powerful neural network to learn structure priors from point cloud, our method can achieve efficient and effective NBV planning. Experiments on multiple datasets show the proposed method outperforms state-of-the-art NBV methods, giving better views for robot system with much less inference time. Furthermore, we demonstrate the robustness of our method against noise and the ability to extend to multi-view system, making it more applicable for various scenarios. Wang Zhao 0001, Yong-Jin Liu 0001 |
IROS | 3 |
| 2020 | Dirichlet energy of Delaunay meshes and intrinsic Delaunay triangulations
Zipeng Ye, Ran Yi 0002, Wen-Yong Gong, Ying He 0001, Yong-Jin Liu 0001 |
Comput. Aided Des. | 5 |
| 2020 | View planning in robot active vision: A survey of systems, algorithms, and applicationsabstractRapid development of artificial intelligence motivates researchers to expand the capabilities of intelligent and autonomous robots. In many robotic applications, robots are required to make planning decisions based on perceptual information to achieve diverse goals in an efficient and effective way. The planning problem has been investigated in active robot vision, in which a robot analyzes its environment and its own state in order to move sensors to obtain more useful information under certain constraints. View planning, which aims to find the best view sequence for a sensor, is one of the most challenging issues in active robot vision. The quality and efficiency of view planning are critical for many robot systems and are influenced by the nature of their tasks, hardware conditions, scanning states, and planning strategies. In this paper, we first summarize some basic concepts of active robot vision, and then review representative work on systems, algorithms and applications from four perspectives: object reconstruction, scene reconstruction, object recognition, and pose estimation. Finally, some potential directions are outlined for future work. Yu-Hui Wen, Wang Zhao 0001, Yong-Jin Liu 0001 |
Comput. Vis. Media | 4 |
| 2020 | Ranking-Preserving Cross-Source Learning for Image Retargeting Quality AssessmentabstractImage retargeting techniques adjust images into different sizes and have attracted much attention recently. Objective quality assessment (OQA) of image retargeting results is often desired to automatically select the best results. Existing OQA methods train a model using some benchmarks (e.g., RetargetMe), in which subjective scores evaluated by users are provided. Observing that it is challenging even for human subjects to give consistent scores for retargeting results of different source images (diff-source-results), in this paper we propose a learning-based OQA method that trains a General Regression Neural Network (GRNN) model based on relative scores-which preserve the ranking-of retargeting results of the same source image (same-source-results). In particular, we develop a novel training scheme with provable convergence that learns a common base scalar for same-source-results. With this source specific offset, our computed scores not only preserve the ranking of subjective scores for same-source-results, but also provide a reference to compare the diff-source-results. We train and evaluate our GRNN model using human preference data collected in RetargetMe. We further introduce a subjective benchmark to evaluate the generalizability of different OQA methods. Experimental results demonstrate that our method outperforms ten representative OQA methods in ranking prediction and has better generalizability to different datasets. Yong-Jin Liu 0001, Yiheng Han, Zipeng Ye, Yukun Lai |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | General Support-Effective Decomposition for Multi-Directional 3-D PrintingabstractWe present a method for fabricating general models with multi-directional 3-D printing systems by printing different model regions along with different directions. The core of our method is a support-effective volume decomposition algorithm that minimizes the area of the regions with large overhangs. A beam-guided searching algorithm with manufacturing constraints determines the optimal volume decomposition, which is represented by a sequence of clipping planes. While current approaches require manually assembling separate components into a final model, our algorithm allows for directly printing the final model in a single pass. It can also be applied to models with loops and handles. A supplementary algorithm generates special supporting structures for models where supporting structures for large overhangs cannot be eliminated. We verify the effectiveness of our method using two hardware systems: a Cartesian-motion-based system and an angular-motion-based system. A variety of 3-D models have been successfully fabricated on these systems. Chenming Wu, Chengkai Dai, Guoxin Fang, Yong-Jin Liu 0001, Charlie C. L. Wang |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2020 | Poisson Vector Graphics (PVG)abstractThis paper presents Poisson vector graphics (PVG), an extension of the popular diffusion curves (DC), for generating smooth-shaded images. Armed with two new types of primitives, called Poisson curves and Poisson regions, PVG can easily produce photorealistic effects such as specular highlights, core shadows, translucency and halos. Within the PVG framework, the users specify color as the Dirichlet boundary condition of diffusion curves and control tone by offsetting the Laplacian of colors, where both controls are simply done by mouse click and slider dragging. PVG distinguishes itself from other diffusion based vector graphics for 3 unique features: 1) explicit separation of colors and tones, which follows the basic drawing principle and eases editing; 2) native support of seamless cloning in the sense that PCs and PRs can automatically fit into the target background; and 3) allowed intersecting primitives (except for DC-DC intersection) so that users can create layers. Through extensive experiments and a preliminary user study, we demonstrate that PVG is a simple yet powerful authoring tool that can produce photo-realistic vector graphics from scratch. Fei Hou 0001, Qian Sun 0003, Zheng Fang 0008, Yong-Jin Liu 0001, Shi-Min Hu 0001, Hong Qin 0001, Aimin Hao, Ying He 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2019 | SketchGAN: Joint Sketch Completion and Recognition With Generative Adversarial NetworkabstractHand-drawn sketch recognition is a fundamental problem in computer vision, widely used in sketch-based image and video retrieval, editing, and reorganization. Previous methods often assume that a complete sketch is used as input; however, hand-drawn sketches in common application scenarios are often incomplete, which makes sketch recognition a challenging problem. In this paper, we propose SketchGAN, a new generative adversarial network (GAN) based approach that jointly completes and recognizes a sketch, boosting the performance of both tasks. Specifically, we use a cascade Encode-Decoder network to complete the input sketch in an iterative manner, and employ an auxiliary sketch recognition task to recognize the completed sketch. Experiments on the Sketchy database benchmark demonstrate that our joint learning approach achieves competitive sketch completion and recognition performance compared with the state-of-the-art methods. Further experiments using several sketch-based applications also validate the performance of our method. Fang Liu 0035, Xiaoming Deng 0001, Yukun Lai, Yong-Jin Liu 0001, CuiXia Ma, Hongan Wang |
CVPR | 4 |
| 2019 | APDrawingGAN: Generating Artistic Portrait Drawings From Face Photos With Hierarchical GANsabstractSignificant progress has been made with image stylization using deep learning, especially with generative adversarial networks (GANs). However, existing methods fail to produce high quality artistic portrait drawings. Such drawings have a highly abstract style, containing a sparse set of continuous graphical elements such as lines, and so small artifacts are much more exposed than for painting styles. Moreover, artists tend to use different strategies to draw different facial features and the lines drawn are only loosely related to obvious image features. To address these challenges, we propose APDrawingGAN, a novel GAN based architecture that builds upon hierarchical generators and discriminators combining both a global network (for images as a whole) and local networks (for individual facial regions). This allows dedicated drawing strategies to be learned for different facial features. Since artists' drawings may not have lines perfectly aligned with image features, we develop a novel loss to measure similarity between generated and artists' drawings based on distance transforms, leading to improved strokes in portrait drawing. To train APDrawingGAN, we construct an artistic drawing dataset containing high-resolution portrait photos and corresponding professional artistic drawings. Extensive experiments, including a user study, show that APDrawingGAN produces significantly better artistic drawings than state-of-the-art methods. Ran Yi 0002, Yong-Jin Liu 0001, Yukun Lai, Paul L. Rosin |
CVPR | 2 |
| 2019 | Fast Computation of Content-Sensitive Superpixels and Supervoxels Using Q-DistancesabstractState-of-the-art researches model the data of images and videos as low-dimensional manifolds and generate superpixels/supervoxels in a content-sensitive way, which is achieved by computing geodesic centroidal Voronoi tessellation (GCVT) on manifolds. However, computing exact GCVTs is slow due to computationally expensive geodesic distances. In this paper, we propose a much faster queue-based graph distance (called q-distance). Our key idea is that for manifold regions in which q-distances are different from geodesic distances, GCVT is prone to placing more generators in them, and therefore after few iterations, the q-distance-induced tessellation is an exact GCVT. This idea works well in practice and we also prove it theoretically under moderate assumption. Our method is simple and easy to implement. It runs 6-8 times faster than state-of-the-art GCVT computation, and has an optimal approximation ratio O(1) and a linear time complexity O(N) for N-pixel images or N-voxel videos. A thorough evaluation of 31 superpixel methods on five image datasets and 8 supervoxel methods on four video datasets shows that our method consistently achieves the best over-segmentation accuracy. We also demonstrate the advantage of our method on one image and two video applications. Zipeng Ye, Ran Yi 0002, Minjing Yu, Yong-Jin Liu 0001, Ying He 0001 |
ICCV | 4 |
| 2019 | An Adaptive Filter for Deep Learning Networks on Large-Scale Point CloudabstractRecently some pioneering works such as PointNet and Point-Net++ successfully introduce deep learning architectures into point cloud analysis. These novel networks take irregular point cloud (i.e., a set of unordered points) as input, in which each point is represented by (x, y, z) coordinates plus some attributes including color, normal, and other local or global features. Despite of their success on various tasks, the computational cost of these networks become extremely high for large-scale point clouds, e.g., containing hundreds of thousands or millions of points. Instead of uniform down-sampling, in this paper, we propose a simple and novel filter that can efficiently filter any large-scale point cloud into thousands of representative points embedded in a high dimensional feature space, such that without changing the existing deep learning networks, simply using our filter as a preprocess, these existing models can work with large-scale point clouds. Experimental results show that by using our proposed filter, the computational cost (measured by floating-point operation) of PointNet and PointNet++ is reduced 30-60 times and the accuracy of semantic segmentation (measured by mean IoU) on ScanNet dataset is improved 5%-15% averagely. Wang Zhao 0001, Ran Yi 0002, Yong-Jin Liu 0001 |
ICIP | 3 |
| 2019 | Vectorization Based Color Transfer for Portrait Images
Ying He 0001, Fei Hou 0001, Juyong Zhang, Anxiang Zeng, Yong-Jin Liu 0001 |
Comput. Aided Des. | 6 |
| 2019 | DE-Path: A Differential-Evolution-Based Method for Computing Energy-Minimizing Paths on Surfaces
Zipeng Ye, Yong-Jin Liu 0001, Jianmin Zheng, Kai Hormann, Ying He 0001 |
Comput. Aided Des. | 2 |
| 2019 | LineUp: Computing Chain-Based Physical TransformationabstractIn this article, we introduce a novel method that can generate a sequence of physical transformations between 3D models with different shape and topology. Feasible transformations are realized on a chain structure with connected components that are 3D printed. Collision-free motions are computed to transform between different configurations of the 3D printed chain structure. To realize the transformation between different 3D models, we first voxelize these input models into a similar number of voxels. The challenging part of our approach is to generate a simple path—as a chain configuration to connect most voxels. A layer-based algorithm is developed with theoretical guarantee of the existence and the path length. We find that collision-free motion sequence can always be generated when using a straight line as the intermediate configuration of transformation. The effectiveness of our method is demonstrated by both the simulation and the experimental tests taken on 3D printed chains. Minjing Yu, Zipeng Ye, Yong-Jin Liu 0001, Ying He 0001, Charlie C. L. Wang |
ACM Trans. Graph. | 3 |
| 2018 | CartoonGAN: Generative Adversarial Networks for Photo CartoonizationabstractIn this paper, we propose a solution to transforming photos of real-world scenes into cartoon style images, which is valuable and challenging in computer vision and computer graphics. Our solution belongs to learning based methods, which have recently become popular to stylize images in artistic forms such as painting. However, existing methods do not produce satisfactory results for cartoonization, due to the fact that (1) cartoon styles have unique characteristics with high level simplification and abstraction, and (2) cartoon images tend to have clear edges, smooth color shading and relatively simple textures, which exhibit significant challenges for texture-descriptor-based loss functions used in existing methods. In this paper, we propose CartoonGAN, a generative adversarial network (GAN) framework for cartoon stylization. Our method takes unpaired photos and cartoon images for training, which is easy to use. Two novel losses suitable for cartoonization are proposed: (1) a semantic content loss, which is formulated as a sparse regularization in the high-level feature maps of the VGG network to cope with substantial style variation between photos and cartoons, and (2) an edge-promoting adversarial loss for preserving clear edges. We further introduce an initialization phase, to improve the convergence of the network to the target manifold. Our method is also much more efficient to train than existing methods. Experimental results show that our method is able to generate high-quality cartoon images from real-world photos (i.e., following specific artists' styles and with clear edges and smooth shading) and outperforms state-of-the-art methods. Yukun Lai, Yong-Jin Liu 0001 |
CVPR | 3 |
| 2018 | Content-Sensitive Supervoxels via Uniform Tessellations on Video ManifoldsabstractSupervoxels are perceptually meaningful atomic regions in videos, obtained by grouping voxels that exhibit coherence in both appearance and motion. In this paper, we propose content-sensitive supervoxels (CSS), which are regularly-shaped 3D primitive volumes that possess the following characteristic: they are typically larger and longer in content-sparse regions (i.e., with homogeneous appearance and motion), and smaller and shorter in content-dense regions (i.e., with high variation of appearance and/or motion). To compute CSS, we map a video Ξ to a 3-dimensional manifold M embedded in R6, whose volume elements give a good measure of the content density in Ξ. We propose an efficient Lloyd-like method with a splitting-merging scheme to compute a uniform tessellation on M, which induces the CSS in Ξ. Theoretically our method has a good competitive ratio O(1). We also present a simple extension of CSS to stream CSS for processing long videos that cannot be loaded into main memory at once. We evaluate CSS, stream CSS and seven representative supervoxel methods on four video datasets. The results show that our method outperforms existing supervoxel methods. Ran Yi 0002, Yong-Jin Liu 0001, Yukun Lai |
CVPR | 2 |
| 2018 | CFD: A Collaborative Feature Difference Method for Spontaneous Micro-Expression SpottingabstractMicro-expression (ME) is a special type of human expression which can reveal the real emotion that people want to conceal. Spontaneous ME (SME) spotting is to identify the subsequences containing SMEs from a long facial video. The study of SME spotting has a significant importance, but is also very challenging due to the fact that in real-world scenarios, SMEs may occur along with normal facial expressions and other prominent motions such as head movements. In this paper, we improve a state-of-the-art SME spotting method called feature difference analysis (FD) in the following two aspects. First, FD relies on a partitioning of facial area into uniform regions of interest (ROIs) and computing features of a selected sequence. We propose a novel evaluation method by utilizing the Fisher linear discriminant to assign a weight for each ROI, leading to more semantically meaningful ROIs. Second, FD only considers two features (LBP and HOOF) independently. We introduce a state-of-the-art MDMO feature into FD and propose a simple yet efficient collaborative strategy to work with two complementary features, i.e., LBP characterizing texture information and MDMO characterizing motion information. We call our improved FD method collaborative feature difference (CFD). Experimental results on two well-established SME datasets SMIC-E and CASME II show that CFD significantly improves the performance of the original FD. Yiheng Han, Bing-Jun Li, Yukun Lai, Yong-Jin Liu 0001 |
ICIP | 4 |
| 2018 | Evaluation on the Compactness of SupervoxelsabstractSupervoxels are perceptually meaningful atomic spatiotemporal regions in videos, which has great potential to reduce the computational complexity of downstream video applications. Many methods have been proposed for generating supervoxels. To effectively evaluate these methods, a novel supervoxel library and benchmark called LIBSVX with seven collected metrics was recently established. In this paper, we propose a new compactness metric which measures the shape regularity of supervoxels and is served as a necessary complement to the existing metrics. To demonstrate its necessity, we first explore the relations between the new metric and existing ones. Correlation analysis shows that the new metric has a weak correlation with (i.e., nearly independent of) existing metrics, and so reflects a new characteristic of supervoxel quality. Second, we investigate two real-world video applications. Experimental results show that the new metric can effectively predict some important application performance, while most existing metrics cannot do so. Ran Yi 0002, Yong-Jin Liu 0001, Yukun Lai |
ICIP | 2 |
| 2018 | Decorating 3D models with Poisson vector graphicsabstractThis paper proposes a novel method for decorating 3D surfaces using a new type of vector graphics, called Poisson Vector Graphics (PVG). Unlike other existing techniques that frequently require local/global parameterization, our approach advocates a parameterization-free paradigm, affording decoration of geometric models with any topological type while minimizing the overall computational expenses. Since PVG supports a set of simple discrete curves, it is straightforward for users to edit colors and synthesize geometry details. Meanwhile, the details could be organized by Poisson Region (PR), leading to much smoother decoration than those of Diffusion Curve (DC). Consequently, it is an ideal tool to create smooth relief. It may be noted that, DC is adequate to create sharp or discontinuous results. But PR is superior to DC, supporting level-of-details editing on meshes thanks to its smoothness. To render PVG on meshes efficiently, we develop a Poisson solver based on harmonic B-splines, which could be constructed using geodesic Voronoi diagram . Our Poisson solver is a local solver for rendering with more flexibility and versatility. We demonstrate the efficacy of our approach on synthetic and real-world 3D models. Fei Hou 0001, Qian Sun 0003, Shi-Qing Xin, Yong-Jin Liu 0001, Wencheng Wang 0001, Hong Qin 0001, Ying He 0001 |
Comput. Aided Des. | 5 |
| 2018 | Micro-expression recognition with small sample size by transferring long-term convolutional neural network
Bing-Jun Li, Yong-Jin Liu 0001, Wen-Jing Yan, Xinyu Ou, Xiaohua Huang 0003, Xiaolan Fu |
Neurocomputing | 3 |
| 2018 | Intrinsic Manifold SLIC: A Simple and Efficient Method for Computing Content-Sensitive SuperpixelsabstractSuperpixels are perceptually meaningful atomic regions that can effectively capture image features. Among various methods for computing uniform superpixels, simple linear iterative clustering (SLIC) is popular due to its simplicity and high performance. In this paper, we extend SLIC to compute content-sensitive superpixels, i.e., small superpixels in content-dense regions with high intensity or colour variation and large superpixels in content-sparse regions. Rather than using the conventional SLIC method that clusters pixels in , we map the input image to a 2-dimensional manifold , whose area elements are a good measure of the content density in . We propose a simple method, called intrinsic manifold SLIC (IMSLIC), for computing a geodesic centroidal Voronoi tessellation (GCVT)-a uniform tessellation-on , which induces the content-sensitive superpixels in . In contrast to the existing algorithms, IMSLIC characterizes the content sensitivity by measuring areas of Voronoi cells on . Using a simple and fast approximation to a closed-form solution, the method can compute the GCVT at a very low cost and guarantees that all Voronoi cells are simply connected. We thoroughly evaluate IMSLIC and compare it with eleven representative methods on the BSDS500 dataset and seven representative methods on the NYUV2 dataset. Computational results show that IMSLIC outperforms existing methods in terms of commonly used quality measures pertaining to superpixels such as compactness, adherence to boundaries, and achievable segmentation accuracy. We also evaluate IMSLIC and seven representative methods in an image contour closure application, and the results on two datasets, WHD and WSD, show that IMSLIC achieves the best foreground segmentation performance. Yong-Jin Liu 0001, Minjing Yu, Bing-Jun Li, Ying He 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Real-Time Movie-Induced Discrete Emotion Recognition from EEG SignalsabstractRecognition of a human's continuous emotional states in real time plays an important role in machine emotional intelligence and human-machine interaction. Existing real-time emotion recognition systems use stimuli with low ecological validity (e.g., picture, sound) to elicit emotions and to recognise only valence and arousal. To overcome these limitations, in this paper, we construct a standardised database of 16 emotional film clips that were selected from over one thousand film excerpts. Based on emotional categories that are induced by these film clips, we propose a real-time movie-induced emotion recognition system for identifying an individual's emotional states through the analysis of brain waves. Thirty participants took part in this study and watched 16 standardised film clips that characterise real-life emotional experiences and target seven discrete emotions and neutrality. Our system uses a 2-s window and a 50 percent overlap between two consecutive windows to segment the EEG signals. Emotional states, including not only the valence and arousal dimensions but also similar discrete emotions in the valence-arousal coordinate space, are predicted in each window. Our real-time system achieves an overall accuracy of 92.26 percent in recognising high-arousal and valenced emotions from neutrality and 86.63 percent in recognising positive from negative emotions. Moreover, our system classifies three positive emotions (joy, amusement, tenderness) with an average of 86.43 percent accuracy and four negative emotions (anger, disgust, fear, sadness) with an average of 65.09 percent accuracy. These results demonstrate the advantage over the existing state-of-the-art real-time emotion recognition systems from EEG signals in terms of classification accuracy and the ability to recognise similar discrete emotions that are close in the valence-arousal coordinate space. Yong-Jin Liu 0001, Minjing Yu, Guozhen Zhao, Jinjing Song, Yan Ge 0007, Yuanchun Shi |
IEEE Trans. Affect. Comput. | 1 |
| 2018 | Delta DLP 3-D Printing of Large ModelsabstractThis paper presents a 3-D printing system that uses a low-cost off-the-shelf consumer projector to fabricate large models. Compared with traditional digital light processing (DLP) 3-D printers using a single vertical carriage, the platform of our DLP 3-D printer using delta mechanism can also move horizontally in the plane. We show that this system can print 3-D models much larger than traditional DLP 3-D printers. The major challenge to realize 3-D printing of large models in our system comes from how to cover a planar polygonal domain by a minimum number of rectangles with fixed size, which is NP-hard. We propose a simple yet efficient approximation algorithm to solve this problem. The key idea is to segment a polygonal domain using its medial axis and afterward merge small parts in the segmentation. Given an arbitrary polygon Q with n generators (i.e., line segments and reflex vertices in Q), we show that the time complexity of our algorithm is O(n2log2n) and the number of output rectangles covering Q is O(Kn), where K is an input-polygon-dependent constant. A physical prototype system is built and several large 3-D models with complex geometric structures have been printed as examples to demonstrate the effectiveness of our approach. Ran Yi 0002, Chenming Wu, Yong-Jin Liu 0001, Ying He 0001, Charlie C. L. Wang |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2018 | Real-Time Assessment of the Cross-Task Mental Workload Using Physiological Measures During Anomaly DetectionabstractThe ability to detect anomalies in perceived stimuli is critical to a broad range of practical and applied activities involving human operators. In this paper, we propose a real-time physiological-based system to assess the cross-task mental workload during anomaly detection. Forty participants were recruited to detect anomalous images from a set of different distracting images (Task I) and abnormal activities from surveillance videos (Task II). In Task I, the task difficulty levels were manipulated by changing the number of anomalies/distracting stimuli (15, 21, 28, or 36) with and without time constraints (i.e., 4 × 2 = 8 task difficulty levels). Physiological and behavioral data from four task difficulty levels were divided into four categories according to subjective ratings of the mental workload. The support vector machine (SVM) classifiers were trained on these data to predict the mental workload categories of: 1) the same four task difficulty levels (within level); and 2) the other four task difficulty levels in Task I (cross level). Within-level classifications (with an average of 95.29%) were more accurate than cross-level classifications (average of 72.2%), which were much more accurate than random level classifications (25%). In Task II, the same participants monitored one, two, or four video clips simultaneously in accordance with three task difficulty levels. The same physiological signals were processed for real-time recognition of a participant's mental workload after he or she completed each activity detection task. The three-class SVM classifiers were trained on physiological data from Task I to predict the mental workload categories of the Task II (cross task), achieving an overall classification accuracy of 53.83%, compared to a 33.33% accuracy at random. These results are discussed in terms of their implications for developing situation-aware recognition systems of the mental workload and adaptive human-computer interaction platforms. Guozhen Zhao, Yong-Jin Liu 0001, Yuanchun Shi |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2018 | Support-free volume printing by multi-axis motionabstractThis paper presents a new method to fabricate 3D models on a robotic printing system equipped with multi-axis motion. Materials are accumulated inside the volume along curved tool-paths so that the need of supporting structures can be tremendously reduced - if not completely abandoned - on all models. Our strategy to tackle the challenge of tool-path planning for multi-axis 3D printing is to perform two successive decompositions, first volume-to-surfaces and then surfaces-to-curves. The volume-to-surfaces decomposition is achieved by optimizing a scalar field within the volume that represents the fabrication sequence. The field is constrained such that its iso-values represent curved layers that are supported from below, and present a convex surface affording for collision-free navigation of the printer head. After extracting all curved layers, the surfaces-to-curves decomposition covers them with tool-paths while taking into account constraints from the robotic printing system. Our method successfully generates tool-paths for 3D printing models with large overhangs and high-genus topology. We fabricated several challenging cases on our robotic platform to verify and demonstrate its capabilities. Chengkai Dai, Charlie C. L. Wang, Chenming Wu, Sylvain Lefebvre 0001, Guoxin Fang, Yong-Jin Liu 0001 |
ACM Trans. Graph. | 6 |
| 2018 | Delaunay mesh simplification with differential evolutionabstractDelaunay meshes (DM) are a special type of manifold triangle meshes --- where the local Delaunay condition holds everywhere --- and find important applications in digital geometry processing. This paper addresses the general DM simplification problem: given an arbitrary manifold triangle mesh M with n vertices and the user-specified resolution m (< n ), compute a Delaunay mesh M * with m vertices that has the least Hausdorffdistance to M. To solve the problem, we abstract the simplification process using a 2D Cartesian grid model, in which each grid point corresponds to triangle meshes with a certain number of vertices and a simplification process is a monotonic path on the grid. We develop a novel differential-evolution-based method to compute a low-cost path, which leads to a high quality Delaunay mesh. Extensive evaluation shows that our method consistently outperforms the existing methods in terms of approximation error. In particular, our method is highly effective for small-scale CAD models and man-made objects with sharp features but less details. Moreover, our method is fully automatic and can preserve sharp features well and deal with models with multiple components, whereas the existing methods often fail. Ran Yi 0002, Yong-Jin Liu 0001, Ying He 0001 |
ACM Trans. Graph. | 2 |
| 2018 | Support-Free HollowingabstractOffsetting-based hollowing is a solid modeling operation widely used in 3D printing, which can change the model's physical properties and reduce the weight by generating voids inside a model. However, a hollowing operation can lead to additional supporting structures for fabrication in interior voids, which cannot be removed. As a consequence, the result of a hollowing operation is affected by these additional supporting structures when applying the operation to optimize physical properties of different models. This paper proposes a support-free hollowing framework to overcome the difficulty of fabricating voids inside a solid. The challenge of computing a support-free hollowing is decomposed into a sequence of shape optimization steps, which are repeatedly applied to interior mesh surfaces. The optimization of physical properties in different applications can be easily integrated into our framework. Comparing to prior approaches that can generate support-free inner structures, our hollowing operation can reduce more volume of material and thus provide a larger solution space for physical optimization. Experimental tests are taken on a number of 3D models to demonstrate the effectiveness of this framework. Weiming Wang 0003, Yong-Jin Liu 0001, Jun Wu 0005, Shengjing Tian, Charlie C. L. Wang, Ligang Liu 0001, Xiuping Liu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2017 | Learning to Rank Retargeted ImagesabstractImage retargeting techniques that adjust images into different sizes have attracted much attention recently. Objective quality assessment (OQA) of image retargeting results is often desired to automatically select the best results. Existing OQA methods output an absolute score for each retargeted image and use these scores to compare different results. Observing that it is challenging even for human subjects to give consistent scores for retargeting results of different source images, in this paper we propose a learning-based OQA method that predicts the ranking of a set of retargeted images with the same source image. We show that this more manageable task helps achieve more consistent prediction to human preference and is sufficient for most application scenarios. To compute the ranking, we propose a simple yet efficient machine learning framework that uses a General Regression Neural Network (GRNN) to model a combination of seven elaborate OQA metrics. We then propose a simple scheme to transform the relative scores output from GRNN into a global ranking. We train our GRNN model using human preference data collected in the elaborate RetargetMe benchmark and evaluate our method based on the subjective study in RetargetMe. Moreover, we introduce a further subjective benchmark to evaluate the generalizability of different OQA methods. Experimental results demonstrate that our method outperforms eight representative OQA methods in ranking prediction and has better generalizability to different datasets. Yong-Jin Liu 0001, Yukun Lai |
CVPR | 2 |
| 2017 | Transforming photos to comics using convolutional neural networksabstractIn this paper, inspired by Gatys's recent work, we propose a novel approach that transforms photos to comics using deep convolutional neural networks (CNNs). While Gatys's method that uses a pre-trained VGG network generally works well for transferring artistic styles such as painting from a style image to a content image, for more minimalist styles such as comics, the method often fails to produce satisfactory results. To address this, we further introduce a dedicated comic style CNN, which is trained for classifying comic images and photos. This new network is effective in capturing various comic styles and thus helps to produce better comic stylization results. Even with a grayscale style image, Gatys's method can still produce colored output, which is not desirable for comics. We develop a modified optimization framework such that a grayscale image is guaranteed to be synthesized. To avoid converging to poor local minima, we further initialize the output image using grayscale version of the content image. Various examples show that our method synthesizes better comic images than the state-of-the-art method. Yukun Lai, Yong-Jin Liu 0001 |
ICIP | 3 |
| 2017 | RoboFDM: A robotic system for support-free fabrication using FDMabstractThis paper presents a robotic system - RoboFDM that targets at printing 3D models without support-structures, which is considered as the major restriction to the flexibility of 3D printing. The hardware of RoboFDM consists of a robotic arm providing 6-DOF motion to the platform of material accumulation and an extruder forming molten filaments of polylactic acid (PLA). The fabrication of 3D models in this system follows the principle of fused decomposition modeling (FDM). Different from conventional FDM, an input model fabricated by RoboFDM is printed along different directions at different places. A new algorithm is developed to decompose models into support-free parts that can be printed one by one in a collision-free sequence. The printing directions of all parts are also determined during the computation of model decomposition. Experiments have been successfully taken on our RoboFDM system to print general freeform objects in a support-free manner. Chenming Wu, Chengkai Dai, Guoxin Fang, Yong-Jin Liu 0001, Charlie C. L. Wang |
ICRA | 4 |
| 2017 | Space complexity of exact discrete geodesic algorithms on regular triangulations
Yong-Jin Liu 0001, Chunxu Xu, Ying He 0001 |
Inf. Process. Lett. | 1 |
| 2017 | Constructing Intrinsic Delaunay Triangulations from the Dual of Geodesic Voronoi DiagramsabstractIntrinsic Delaunay triangulation (IDT) naturally generalizes Delaunay triangulation from R 2 to curved surfaces. Due to many favorable properties, the IDT whose vertex set includes all mesh vertices is of particular interest in polygonal mesh processing. To date, the only way for constructing such IDT is the edge-flipping algorithm, which iteratively flips non-Delaunay edges to become locally Delaunay. Although this algorithm is conceptually simple and guarantees to terminate in finite steps, it has no known time complexity and may also produce triangulations containing faces with only two edges. This article develops a new method to obtain proper IDTs on manifold triangle meshes. We first compute a geodesic Voronoi diagram (GVD) by taking all mesh vertices as generators and then find its dual graph. The sufficient condition for the dual graph to be a proper triangulation is that all Voronoi cells satisfy the so-called closed ball property. To guarantee the closed ball property everywhere, a certain sampling criterion is required. For Voronoi cells that violate the closed ball property, we fix them by computing topologically safe regions, in which auxiliary sites can be added without changing the topology of the Voronoi diagram beyond them. Given a mesh with n vertices, we prove that by adding at most O ( n ) auxiliary sites, the computed GVD satisfies the closed ball property, and hence its dual graph is a proper IDT. Our method has a theoretical worst-case time complexity O ( n 2 + tn log n ), where t is the number of obtuse angles in the mesh. Computational results show that it empirically runs in linear time on real-world models. Yong-Jin Liu 0001, Chunxu Xu, Ying He 0001 |
ACM Trans. Graph. | 1 |
| 2017 | Objective Quality Prediction of Image Retargeting AlgorithmsabstractQuality assessment of image retargeting results is useful when comparing different methods. However, performing the necessary user studies is a long, cumbersome process. In this paper, we propose a simple yet efficient objective quality assessment method based on five key factors: i) preservation of salient regions; ii) analysis of the influence of artifacts; iii) preservation of the global structure of the image; iv) compliance with well-established aesthetics rules; and v) preservation of symmetry. Experiments on the RetargetMe benchmark, as well as a comprehensive additional user study, demonstrate that our proposed objective quality assessment method outperforms other existing metrics, while correlating better with human judgements. This makes our metric a good predictor of subjective preference. Yun Liang 0003, Yong-Jin Liu 0001, Diego Gutierrez |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2017 | Cross section-based hollowing and structural enhancement
Weiming Wang 0003, Baojun Li, Sicheng Qian, Yong-Jin Liu 0001, Charlie C. L. Wang, Ligang Liu 0001, Xiuping Liu |
Vis. Comput. | 4 |
| 2016 | Manifold SLIC: A Fast Method to Compute Content-Sensitive SuperpixelsabstractSuperpixels are perceptually meaningful atomic regions that can effectively capture image features. Among various methods for computing uniform superpixels, simple linear iterative clustering (SLIC) is popular due to its simplicity and high performance. In this paper, we extend SLIC to compute content-sensitive superpixels, i.e., small superpixels in content-dense regions (e.g., with high intensity or color variation) and large superpixels in content-sparse regions. Rather than the conventional SLIC method that clusters pixels in ℝ5, we map the image I to a 2-dimensional manifold M ⊂ ℝ5, whose area elements are a good measure of the content density in I. We propose an efficient method to compute restricted centroidal Voronoi tessellation (RCVT) - a uniform tessellation - on M, which induces the content-sensitive superpixels in I. Unlike other algorithms that characterize content-sensitivity by geodesic distances, manifold SLIC tackles the problem by measuring areas of Voronoi cells on M, which can be computed at a very low cost. As a result, it runs 10 times faster than the state-of-the-art content-sensitive superpixels algorithm. We evaluate manifold SLIC and seven representative methods on the BSDS500 benchmark and observe that our method outperforms the existing methods. Yong-Jin Liu 0001, Cheng-Chi Yu, Minjing Yu, Ying He 0001 |
CVPR | 1 |
| 2016 | Delta DLP 3D printing with large sizeabstractWe present a delta DLP 3D printer with large size in this paper. Compared with traditional DLP 3D printers that use a low-cost off-the-shelf consumer projector and a single vertical carriage, the platform of our delta DLP 3D printer can also move horizontally in the plane. We show that this structure allows the printer to have a larger printing area than the projection area of a projector. Our system can print 3D models much larger than traditional DLP 3D printers. The major challenge to realize delta 3D printing with large size comes from how to partition an arbitrary planar polygonal shape (possibly with holes or multiple disjoint polygons) into a minimum number of rectangles with fixed size, which is NP-hard. We propose a simple yet efficient approximation algorithm to solve this problem. The time complexity of our algorithm is O(n3log n), where n is the number of edges in the polygonal shape. A physical prototype system is built and several large 3D models with complex geometric structures have been printed as examples to demonstrate the effectiveness of our approach. Chenming Wu, Ran Yi 0002, Yong-Jin Liu 0001, Ying He 0001, Charlie C. L. Wang |
IROS | 3 |
| 2016 | Barehanded music: real-time hand interaction for virtual pianoabstractThis paper presents an efficient data-driven approach to track fingertip and detect finger tapping for virtual piano using an RGB-D camera. We collect 7200 depth images covering the most common finger articulation for playing piano, and train a random regression forest using depth context features of randomly sampled pixels in training images. In the online tracking stage, we firstly segment the hand from the plane in contact by fusing the information from both color and depth images. Then we use the trained random forest to estimate the 3D position of fingertips and wrist in each frame, and predict finger tapping based on the estimated fingertip motion. Finally, we build a kinematic chain and recover the articulation parameters for each finger. In contrast to the existing hand tracking algorithms that often require hands are in the air and cannot interact with physical objects, our method is designed for hand interaction with planar objects, which is desired for the virtual piano application. Using our prototype system, users can put their hands on a desk, move them sideways and then tap fingers on the desk, like playing a real piano. Preliminary results show that our method can recognize most of the beginner's piano-playing gestures in realtime for soothing rhythms. Hui Liang 0003, Jin Wang 0018, Qian Sun 0003, Yong-Jin Liu 0001, Junsong Yuan 0001, Jun Luo 0001, Ying He 0001 |
I3D | 4 |
| 2016 | Solving the initial value problem of discrete geodesics
Peng Cheng 0008, Chunyan Miao, Yong-Jin Liu 0001, Changhe Tu, Ying He 0001 |
Comput. Aided Des. | 3 |
| 2016 | Cognitive mechanism related to line drawings and its applications in intelligent process of visual media: a survey
Yong-Jin Liu 0001, Minjing Yu, Qiu-Fang Fu, Ye Liu 0010, Lexing Xie |
Frontiers Comput. Sci. | 1 |
| 2016 | A PMJ-inspired cognitive framework for natural scene categorization in line drawings
Minjing Yu, Yong-Jin Liu 0001, Qiu-Fang Fu, Xiaolan Fu |
Neurocomputing | 2 |
| 2016 | A Main Directional Mean Optical Flow Feature for Spontaneous Micro-Expression RecognitionabstractMicro-expressions are brief facial movements characterized by short duration, involuntariness and low intensity. Recognition of spontaneous facial micro-expressions is a great challenge. In this paper, we propose a simple yet effective Main Directional Mean Optical-flow (MDMO) feature for micro-expression recognition. We apply a robust optical flow method on micro-expression video clips and partition the facial area into regions of interest (ROIs) based partially on action units. The MDMO is a ROI-based, normalized statistic feature that considers both local statistic motion information and its spatial location. One of the significant characteristics of MDMO is that its feature dimension is small. The length of a MDMO feature vector is 36 × 2 = 72, where 36 is the number of ROIs. Furthermore, to reduce the influence of noise due to head movements, we propose an optical-flow-driven method to align all frames of a micro-expression video clip. Finally, a SVM classifier with the proposed MDMO feature is adopted for micro-expression recognition. Experimental results on three spontaneous micro-expression databases, namely SMIC, CASME and CASME II, show that the MDMO can achieve better performance than two state-of-the-art baseline features, i.e., LBP-TOP and HOOF. Yong-Jin Liu 0001, Jinkai Zhang, Wen-Jing Yan, Guoying Zhao 0001, Xiaolan Fu |
IEEE Trans. Affect. Comput. | 1 |
| 2016 | An Interactive SpiralTape Video SummarizationabstractA majority of video summarization systems use linear representations, such as rectangular storyboards and timelines at linear scales. In this paper, we propose a novel nonlinear dynamic representation called SpiralTape that summarizes a video in a smooth spiral pattern. SpiralTape provides an unusual and fresh activity suitable for stimulating environments such as science and technology museums, in which children or young individuals can have enjoyable experiences that create meaningful learning outcomes. In addition, SpiralTape provides an uninterrupted overall structure of video content and takes design principles including compactness, continuity, efficient overview, and interactivity into consideration. A working SpiralTape system was developed and deployed in pilot applications and exhibitions. Elaborate user studies with evaluation benchmarks on multiple metrics were conducted to compare SpiralTape with two representative linear video summarization methods and a state-of-the-art radial video visualization. The evaluation results demonstrate the effectiveness and natural interaction performance of SpiralTape. Yong-Jin Liu 0001, CuiXia Ma, Guozhen Zhao, Xiaolan Fu, Hongan Wang, Guozhong Dai, Lexing Xie |
IEEE Trans. Multim. | 1 |
| 2016 | Visualizing and Analyzing Video Content With Interactive Scalable MapsabstractVisualizing and communicating insights through maps offers an intuitive and familiar way to explore large-scale dynamic relational data. In this paper, we present VideoMap, which is a novel approach for presenting and interacting with relational video content by taking advantage of the map metaphor. VideoMap employs a metaphor to visualize video content by elements of a map with the aim of enabling exploration of video content as if reading a map. Video content is visualized in a hierarchal structure from a very large scale to a small scale of finely detailed representation. VideoMap recognizes a small set of sketch gestures for semantic zooming in and out, annotating the map, and automatically completing path navigation. To achieve this, VideoMap synthesizes map-derived visuals and binds them to the underlying data by operating the map with sketch interaction to facilitate interactive exploration. Extensive user studies were conducted to evaluate VideoMap, and the results demonstrated the effectiveness of VideoMap for facilitating the exploration and understanding of large video content. CuiXia Ma, Yong-Jin Liu 0001, Guozhen Zhao, Hongan Wang |
IEEE Trans. Multim. | 2 |
| 2016 | Manifold differential evolution (MDE): a global optimization method for geodesic centroidal voronoi tessellations on meshesabstractComputing centroidal Voronoi tessellations (CVT) has many applications in computer graphics. The existing methods, such as the Lloyd algorithm and the quasi-Newton solver, are efficient and easy to implement; however, they compute only the local optimal solutions due to the highly non-linear nature of the CVT energy. This paper presents a novel method, called manifold differential evolution (MDE), for computing globally optimal geodesic CVT energy on triangle meshes. Formulating the mutation operator using discrete geodesics, MDE naturally extends the powerful differential evolution framework from Euclidean spaces to manifold domains. Under mild assumptions, we show that MDE has a provable probabilistic convergence to the global optimum. Experiments on a wide range of 3D models show that MDE consistently out-performs the existing methods by producing results with lower energy. Thanks to its intrinsic and global nature, MDE is insensitive to initialization and mesh tessellation. Moreover, it is able to handle multiply-connected Voronoi cells, which are challenging to the existing geodesic CVT methods. Yong-Jin Liu 0001, Chunxu Xu, Ran Yi 0002, Ying He 0001 |
ACM Trans. Graph. | 1 |
| 2016 | Styling Evolution for Tight-Fitting GarmentsabstractWe present an evolution method for designing the styling curves of garments. The procedure of evolution is driven by aesthetics-inspired scores to evaluate the quality of styling designs, where the aesthetic considerations are represented in the form of streamlines on human bodies. A dual representation is introduced in our platform to process the styling curves of designs, based on which robust methods for realizing the operations of evolution are developed. Starting from a given set of styling designs on human bodies, we demonstrate the effectiveness of set evolution inspired by aesthetic factors. The evolution is adaptive to the change of aesthetic inspirations. By this adaptation, our platform can automatically generate new designs fulfilling the demands of variations in different human bodies and poses. Tsz-Ho Kwok, Yanqiu Zhang, Charlie C. L. Wang, Yong-Jin Liu 0001, Kai Tang 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2016 | A Robust Divide and Conquer Algorithm for Progressive Medial Axes of Planar ShapesabstractThe medial axis is an important shape representation that finds a wide range of applications in shape analysis. For large-scale shapes of high resolution, a progressive medial axis representation that starts with the lowest resolution and gradually adds more details is desired. In this paper, we propose a fast and robust geometric algorithm that computes progressive medial axes of a large-scale planar shape. The key ingredient of our method is a novel structural analysis of merging medial axes of two planar shapes along a shared boundary. Our method is robust by separating the analysis of topological structure from numerical computation. Our method is also fast and we show that the time complexity of merging two medial axes is$O(n\;\log n_v)$, where$n$is the number of total boundary generators,$n_v$is strictly smaller than$n$and behaves as a small constant in all our experiments. Experiments on large-scale polygonal data and comparison with state-of-the-art methods show the efficiency and effectiveness of the proposed method. Yong-Jin Liu 0001, Cheng-Chi Yu, Minjing Yu, Kai Tang 0001, Deok-Soo Kim |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2015 | Intrinsic computation of centroidal Voronoi tessellation (CVT) on meshes
Xiang Ying, Yong-Jin Liu 0001, Shi-Qing Xin, Wenping Wang 0001, Xianfeng Gu, Wolfgang Müller-Wittig, Ying He 0001 |
Comput. Aided Des. | 3 |
| 2015 | A unified framework for isotropic meshing based on narrow-band Euclidean distance transformationabstractIn this paper, we propose a simple-yet-effective method for isotropic meshing relying on Euclidean distance transformation based centroidal Voronoi tessellation (CVT). Our approach improves the performance and robustness of computing CVT on curved domains while simultaneously providing high-quality output meshes. While conventional extrinsic methods compute CVTs in the entire volume bounded by the input model, we restrict the computation to a 3D shell of user-controlled thickness. Taking voxels which contain surface samples as sites, we compute the exact Euclidean distance transform on the GPU. Our algorithm is parallel and memory-efficient, and can construct the shell space for resolutions up to 2048 3 at interactive speed. The 3D centroidal Voronoi tessellation and restricted Voronoi diagrams are also computed efficiently on the GPU. Since the shell space can bridge holes and gaps smaller than a certain tolerance, and tolerate non-manifold edges and degenerate triangles, our algorithm can handle models with such defects, which typically cause conventional remeshing methods to fail. Our method can process implicit surfaces, polyhedral surfaces, and point clouds in a unified framework. Computational results show that our GPU-based isotropic meshing algorithm produces results comparable to state-of- the-art techniques, but is significantly faster than conventional CPU-based implementations. Yuen-Shan Leung, Ying He 0001, Yong-Jin Liu 0001, Charlie C. L. Wang |
Comput. Vis. Media | 4 |
| 2015 | Semi-Continuity of Skeletons in Two-Manifold and Discrete Voronoi ApproximationabstractThe skeleton of a 2D shape is an important geometric structure in pattern analysis and computer vision. In this paper we study the skeleton of a 2D shape in a two-manifold$\mathcal {M}$, based on a geodesic metric. We present a formal definition of the skeleton$S(\Omega )$for a shape$\Omega$in$\mathcal {M}$and show several properties that make$S(\Omega )$distinct from its Euclidean counterpart in$\mathbb {R}^2$. We further prove that for a shape sequence$\lbrace \Omega _i\rbrace$that converge to a shape$\Omega$in$\mathcal {M}$, the mapping$\Omega \rightarrow \overline{S}(\Omega )$is lower semi-continuous. A direct application of this result is that we can use a set$P$of sample points to approximate the boundary of a 2D shape$\Omega$in$\mathcal {M}$, and the Voronoi diagram of$P$inside$\Omega \subset \mathcal {M}$gives a good approximation to the skeleton$S(\Omega )$. Examples of skeleton computation in topography and brain morphometry are illustrated. Yong-Jin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Efficient construction and simplification of Delaunay meshesabstractDelaunay meshes (DM) are a special type of triangle mesh where the local Delaunay condition holds everywhere. We present an efficient algorithm to convert an arbitrary manifold triangle mesh M into a Delaunay mesh. We show that the constructed DM has O ( Kn ) vertices, where n is the number of vertices in M and K is a model-dependent constant. We also develop a novel algorithm to simplify Delaunay meshes, allowing a smooth choice of detail levels. Our methods are conceptually simple, theoretically sound and easy to implement. The DM construction algorithm also scales well due to its O ( nK log K ) time complexity. Delaunay meshes have many favorable geometric and numerical properties. For example, a DM has exactly the same geometry as the input mesh, and it can be encoded by any mesh data structure. Moreover, the empty geodesic circumcircle property implies that the commonly used cotangent Laplace-Beltrami operator has non-negative weights. Therefore, the existing digital geometry processing algorithms can benefit the numerical stability of DM without changing any codes. We observe that DMs can improve the accuracy of the heat method for computing geodesic distances. Also, popular parameterization techniques, such as discrete harmonic mapping, produce more stable results on the DMs than on the input meshes. Yong-Jin Liu 0001, Chunxu Xu, Ying He 0001 |
ACM Trans. Graph. | 1 |
| 2015 | Fast Wavefront Propagation (FWP) for Computing Exact Geodesic Distances on MeshesabstractComputing geodesic distances on triangle meshes is a fundamental problem in computational geometry and computer graphics. To date, two notable classes of algorithms, the Mitchell-Mount-Papadimitriou (MMP) algorithm and the Chen-Han (CH) algorithm, have been proposed. Although these algorithms can compute exact geodesic distances if numerical computation is exact, they are computationally expensive, which diminishes their usefulness for large-scale models and/or time-critical applications. In this paper, we propose the fast wavefront propagation (FWP) framework for improving the performance of both the MMP and CH algorithms. Unlike the original algorithms that propagate only a single window (a data structure locally encodes geodesic information) at each iteration, our method organizes windows with a bucket data structure so that it can process a large number of windows simultaneously without compromising wavefront quality. Thanks to its macro nature, the FWP method is less sensitive to mesh triangulation than the MMP and CH algorithms. We evaluate our FWP-based MMP and CH algorithms on a wide range of large-scale real-world models. Computational results show that our method can improve the speed by a factor of 3-10. Chunxu Xu, Tuanfeng Y. Wang, Yong-Jin Liu 0001, Ligang Liu 0001, Ying He 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2014 | Sketch2Jewelry: Semantic feature modeling for sketch-based jewelry design
Long Zeng 0001, Yong-Jin Liu 0001, Jin Wang 0015, Matthew M. F. Yuen |
Comput. Graph. | 2 |
| 2014 | Polyline-sourced Geodesic Voronoi Diagrams on Triangle MeshesabstractAbstract This paper studies the Voronoi diagrams on 2‐manifold meshes based on geodesic metric (a.k.a. geodesic Voronoi diagrams or GVDs), which have polyline generators. We show that our general setting leads to situations more complicated than conventional 2D Euclidean Voronoi diagrams as well as point‐source based GVDs, since a typical bisector contains line segments, hyperbolic segments and parabolic segments. To tackle this challenge, we introduce a new concept, called local Voronoi diagram (LVD), which is a combination of additively weighted Voronoi diagram and line‐segment Voronoi diagram on a mesh triangle. We show that when restricting on a single mesh triangle, the GVD is a subset of the LVD and only two types of mesh triangles can contain GVD edges. Based on these results, we propose an efficient algorithm for constructing the GVD with polyline generators. Our algorithm runs in O(nNlogN) time and takes O(nN) space on an n‐face mesh with m generators, where N = max{m, n}. Computational results on real‐world models demonstrate the efficiency of our algorithm. Chunxu Xu, Yong-Jin Liu 0001, Qian Sun 0003, Jinyan Li 0001, Ying He 0001 |
Comput. Graph. Forum | 2 |
| 2014 | A computational cognition model of perception, memory, and judgment
Xiaolan Fu, Lianhong Cai, Ye Liu 0010, Jia Jia 0001, Zhang Yi 0001, Guozhen Zhao, Yong-Jin Liu 0001, Changxu Wu |
Sci. China Inf. Sci. | 8 |
| 2014 | A global energy optimization framework for 2.1D sketch extraction from monocular images
Cheng-Chi Yu, Yong-Jin Liu 0001, Tianfu Wu 0001, Kai-Yun Li, Xiaolan Fu |
Graph. Model. | 2 |
| 2014 | For micro-expression recognition: Database and suggestions
Wen-Jing Yan, Yong-Jin Liu 0001, Xiaolan Fu |
Neurocomputing | 3 |
| 2014 | A Sketch-Based Approach for Interactive Organization of Video ClipsabstractWith the rapid growth of video resources, techniques for efficient organization of video clips are becoming appealing in the multimedia domain. In this article, a sketch-based approach is proposed to intuitively organize video clips by: (1) enhancing their narrations using sketch annotations and (2) structurizing the organization process by gesture-based free-form sketching on touch devices. There are two main contributions of this work. The first is a sketch graph, a novel representation for the narrative structure of video clips to facilitate content organization. The second is a method to perform context-aware sketch recommendation scalable to large video collections, enabling common users to easily organize sketch annotations. A prototype system integrating the proposed approach was evaluated on the basis of five different aspects concerning its performance and usability. Two sketch searching experiments showed that the proposed context-aware sketch recommendation outperforms, in terms of accuracy and scalability, two state-of-the-art sketch searching methods. Moreover, a user study showed that the sketch graph is consistently preferred over traditional representations such as keywords and keyframes. The second user study showed that the proposed approach is applicable in those scenarios where the video annotator and organizer were the same person. The third user study showed that, for video content organization, using sketch graph users took on average 1/3 less time than using a mass-market tool Movie Maker and took on average 1/4 less time than using a state-of-the-art sketch alternative. These results demonstrated that the proposed sketch graph approach is a promising video organization tool. Yong-Jin Liu 0001, CuiXia Ma, Qiu-Fang Fu, Xiaolan Fu, Sheng Feng Qin, Lexing Xie |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2013 | A distributed computational cognitive model for object recognition
Yong-Jin Liu 0001, Qiu-Fang Fu, Ye Liu 0010, Xiaolan Fu |
Sci. China Inf. Sci. | 1 |
| 2013 | The complexity of geodesic Voronoi diagrams on triangulated 2-manifold surfaces
Yong-Jin Liu 0001, Kai Tang 0001 |
Inf. Process. Lett. | 1 |
| 2013 | Collaborative Interaction for Videos on Mobile Devices Based on Sketch Gestures
Jinkai Zhang, CuiXia Ma, Yong-Jin Liu 0001, Qiu-Fang Fu, Xiaolan Fu |
J. Comput. Sci. Technol. | 3 |
| 2013 | User-Adaptive Sketch-Based 3-D CAD Model Retrievalabstract3-D CAD models are an important digital resource in the manufacturing industry. 3-D CAD model retrieval has become a key technology in product lifecycle management enabling the reuse of existing design data. In this paper, we propose a new method to retrieve 3-D CAD models based on 2-D pen-based sketch inputs. Sketching is a common and convenient method for communicating design intent during early stages of product design, e.g., conceptual design. However, converting sketched information into precise 3-D engineering models is cumbersome, and much of this effort can be avoided by reuse of existing data. To achieve this purpose, we present a user-adaptive sketch-based retrieval method in this paper. The contributions of this work are twofold. First, we propose a statistical measure for CAD model retrieval: the measure is based on sketch similarity and accounts for users' drawing habits. Second, for 3-D CAD models in the database, we propose a sketch generation pipeline that represents each 3-D CAD model by a small yet sufficient set of sketches that are perceptually similar to human drawings. User studies and experiments that demonstrate the effectiveness of the proposed method in the design process are presented. Yong-Jin Liu 0001, Ajay Joneja, CuiXia Ma, Xiaolan Fu, Dawei Song 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2013 | Cylinder Detection in Large-Scale Point Cloud of Pipeline PlantabstractThe huge number of points scanned from pipeline plants make the plant reconstruction very difficult. Traditional cylinder detection methods cannot be applied directly due to the high computational complexity. In this paper, we explore the structural characteristics of point cloud in pipeline plants and define a structure feature. Based on the structure feature, we propose a hierarchical structure detection and decomposition method that reduces the difficult pipeline-plant reconstruction problem in IR³ into a set of simple circle detection problems in IR². Experiments with industrial applications are presented, which demonstrate the efficiency of the proposed structure detection method. Yong-Jin Liu 0001, Ji-Chun Hou, Ji-Cheng Ren, Wei-Qing Tang |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2012 | 2D-Line-Drawing-Based 3D Object Recognition
Yong-Jin Liu 0001, Qiu-Fang Fu, Ye Liu 0010, Xiaolan Fu |
CVM | 1 |
| 2012 | Q--Complex: Efficient non-manifold boundary representation with inclusion topology
Long Zeng 0001, Yong-Jin Liu 0001, Matthew M. F. Yuen |
Comput. Aided Des. | 2 |
| 2012 | Least squares quasi-developable mesh approximation
Long Zeng 0001, Yong-Jin Liu 0001, Matthew M. F. Yuen |
Comput. Aided Geom. Des. | 2 |
| 2012 | Sketch-Based Annotation and Visualization in Video AuthoringabstractAuthoring context-aware, interactive video representation is usually a complex process. A user-friendly multimedia authoring environment is thus solicited to explore and express users' design ideas efficiently and naturally. In this paper we present a sketch-based two-layer representation, called scene structure graph (SSG), to facilitate the video authoring process. One layer in SSG uses sketches as a concise form with which the visualization of scene information is easily understood and the other layer uses a graph to represent and edit the narrative structure in the authoring process. With SSG, the authoring process works in two stages. In the first stage, various sketch forms such as symbols and hand-drawing illustrations are used as basic primitives to annotate the video clips and the hyperlinks encoding spatio-temporal relations are established in SSG. In the second stage, sketches in SSGs are modified and new SSG is composed for any particular authoring purpose. Three user studies are elaborated, showing that the SSG is user-friendly and can achieve a good balance between expressiveness of users' intent and ease of use for authoring of interactive video. CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang, Dongxing Teng, Guozhong Dai |
IEEE Trans. Multim. | 2 |
| 2012 | 3D model retrieval based on color + geometry signatures
Yong-Jin Liu 0001, Yi-Fu Zheng, Yuming Xuan, Xiaolan Fu |
Vis. Comput. | 1 |
| 2011 | Industrial design using interpolatory discrete developable surfaces
Yong-Jin Liu 0001, Kai Tang 0001, Wen-Yong Gong, Tie-Ru Wu |
Comput. Aided Des. | 1 |
| 2011 | Image Retargeting Quality AssessmentabstractAbstract Content‐aware image retargeting is a technique that can flexibly display images with different aspect ratios and simultaneously preserve salient regions in images. Recently many image retargeting techniques have been proposed. To compare image quality by different retargeting methods fast and reliably, an objective metric simulating the human vision system (HVS) is presented in this paper. Different from traditional objective assessment methods that work in bottom‐up manner (i.e., assembling pixel‐level features in a local‐to‐global way), in this paper we propose to use a reverse order (top‐down manner) that organizes image features from global to local viewpoints, leading to a new objective assessment metric for retargeted images. A scale‐space matching method is designed to facilitate extraction of global geometric structures from retargeted images. By traversing the scale space from coarse to fine levels, local pixel correspondence is also established. The objective assessment metric is then based on both global geometric structures and local pixel correspondence. To evaluate color images, CIE L*a*b* color space is utilized. Experimental results are obtained to measure the performance of objective assessments with the proposed metric. The results show good consistency between the proposed objective metric and subjective assessment by human observers. Yong-Jin Liu 0001, Yuming Xuan, Xiaolan Fu |
Comput. Graph. Forum | 1 |
| 2011 | Construction of Iso-Contours, Bisectors, and Voronoi Diagrams on Triangulated SurfacesabstractIn the research of computer vision and machine perception, 3D objects are usually represented by 2-manifold triangular meshes M. In this paper, we present practical and efficient algorithms to construct iso-contours, bisectors, and Voronoi diagrams of point sites on M, based on an exact geodesic metric. Compared to euclidean metric spaces, the Voronoi diagrams on M exhibit many special properties that fail all of the existing euclidean Voronoi algorithms. To provide practical algorithms for constructing geodesic-metric-based Voronoi diagrams on M, this paper studies the analytic structure of iso-contours, bisectors, and Voronoi diagrams on M. After a necessary preprocessing of model M, practical algorithms are proposed for quickly obtaining full information about iso--contours, bisectors, and Voronoi diagrams on M. The complexity of the construction algorithms is also analyzed. Finally, three interesting applications-surface sampling and reconstruction, 3D skeleton extraction, and point pattern analysis-are presented that show the potential power of the proposed algorithms in pattern analysis. Yong-Jin Liu 0001, Zhanqing Chen, Kai Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | KnitSketch: A Sketch Pad for Conceptual Design of 2D Garment PatternsabstractIn this paper, we present a new sketch-based system - KnitSketch, to improve the efficiency of process planning for knitting garments at an early design stage. The KnitSketch system utilizes sketching interface with the pen-paper metaphor and users only need to draw outlines of different parts of the garment. Based on sketching understanding, the system automatically makes reasonable geometric inferences about the process-planning data of the garment. The system is designed for nonprofessional users and can design diverse garment styles by freehand drawings. The contributions of this work include contextual extraction of reusable data from sketches, a MDG structure for sketch beautification, and an integrated system with natural expression and effective communication that reduces the cognitive load of human beings. User experience shows that the proposed system helps designers focus on the task instead of the designing tools, and thus improves the efficiency and productivity of human beings. CuiXia Ma, Yong-Jin Liu 0001, Dongxing Teng, Hongan Wang, Guozhong Dai |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2010 | A Semantic Feature Model in Concurrent EngineeringabstractConcurrent engineering (CE) is a methodology applied to product lifecycle development so that high quality, well designed products can be provided at lower prices and in less time. Many research works have been proposed for efficiently modeling of different domains in CE. However, an integration of these works with consistent data flow is absent and still in great demand in industry. In this paper, we present a generic integration framework with a semantic feature model for knowledge representation and reasoning across domains in CE. An implementation of the proposed semantic feature model is presented to demonstrate its advantage in knowledge representation by feature transformation across domains in CE. Yong-Jin Liu 0001, Kam-Lung Lai, Matthew M. F. Yuen |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2010 | Some notes on maximal arc intersection of spherical polygons: its NP\mathcal{NP} -hardness and approximation algorithms
Yong-Jin Liu 0001, Kai Tang 0001 |
Vis. Comput. | 1 |
| 2009 | A semantic feature language for concurrent engineeringabstractIn this paper, we present a semantic feature language for knowledge representation and reasoning across different domains in concurrent engineering (CE). The proposed language utilizes a feature model which is depicted in two levels. First a unified and formal framework is presented for semantic features in CE. The presented semantic features are general and a feature formal language is used to characterize the engineering data in multi-discipline domains at the most abstract level. At the details level, the general semantic features are realized in a hierarchical manner with the aids of object-oriented programming. A preliminary implementation of the proposed semantic feature model is also presented. Yong-Jin Liu 0001, D. Lai, Matthew M. F. Yuen |
CAD/Graphics | 1 |
| 2009 | On the performance of maximal intersection of spherical polygons by arcsabstractAn important real-world optimization problem in manufacturing industry is to determine optimal workpiece setups for 4-axis NC machining. In this paper we reveal some interesting relations between this optimal workpiece setup problem and the two classic NP-hard problems in complexity theory (i.e, the vertex cover problem and the set cover problem). These relations immediately show the following results. First the optimal workpiece setup problem is NP-hard. Secondly, the greedy algorithm proposed in [Comput. Aided Des. 35 (2003) pp. 1269-1285] for the optimal workpiece setup problem has the performance ratio bounded by O(ln n-ln ln n+0.78), where n is the number of spherical polygons in the ground set. Yong-Jin Liu 0001, Kai Tang 0001 |
CAD/Graphics | 1 |
| 2009 | Stripification of Free-Form Surfaces With Global Error Bounds for Developable ApproximationabstractDevelopable surfaces have many desired properties in the manufacturing process. Since most existing CAD systems utilize tensor-product parametric surfaces including B-splines as design primitives, there is a great demand in industry to convert a general free-form parametric surface within a prescribed global error bound into developable patches. In this paper, we propose a practical and efficient solution to approximate a rectangular parametric surface with a small set ofC0-joint developable strips. The key contribution of the proposed algorithm is that, several optimization problems are elegantly solved in a sequence that offers a controllable global error bound on the developable surface approximation. Experimental results are presented to demonstrate the effectiveness and stability of the proposed algorithm. Yong-Jin Liu 0001, Yukun Lai, Shi-Min Hu 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2008 | Planar Shape Matching and Feature Extraction Using Shape Profile
Yong-Jin Liu 0001, Terry K. K. Chang, Matthew M. F. Yuen |
GMP | 1 |
| 2008 | Note on Industrial Applications of Hu's Surface Extension Algorithm
Yong-Jin Liu 0001, Yukun Lai |
GMP | 2 |
| 2008 | Fairing wireframes in industrial surface designabstractWireframe is a modeling tool widely used in industrial geometric design. The term wireframe refers to two sets of curves, with the property that each curve from one set intersects with each curve from the other set. Akin to the mu-, v-isocurves in a tensor-product surface, the two sets of curves in a wireframe span an underlying surface. In many industrial design activities, wireframes are usually set up and adjusted by the designers before the whole surfaces are reconstructed. For adjustment, the fairness of wireframe has a direct influence on the quality of the underlying surface. Wireframe fairing is significantly different from fairing individual curves in that intersections should be preserved and kept in the same order. In this paper, we first present a technique for wireframe fairing by fixing the parameters during fairing. The limitation of fixed parameters is further released by an iterative gradient descent optimization method with step-size control. Experimental results show that our solution is efficient, and produces reasonably fairing results of the wireframes. Yukun Lai, Yong-Jin Liu 0001, Shi-Min Hu 0001 |
Shape Modeling International | 2 |
| 2008 | Geometry-optimized virtual human head and its applications
Yong-Jin Liu 0001, Matthew M. F. Yuen |
Comput. Graph. | 1 |
| 2007 | A New Canonical Model of Virtual Human HeadabstractIn this paper, a new canonical model of virtual human head is presented. By using a uniform two-step refinement operator, the canonical model has a natural multiresolution structure. In the coarsest level, the canonical model is represented by a control mesh M0which serves as the anatomical structure of the head. To achieve fine geometric details over M0, a multi-level detail set is assigned to the corresponding control vertices in M0, with the aid of a uniquely defined local frame field. Compared to the previously work, the proposed canonical model not only captures the physical characteristics of the head, but also is optimal in the geometric sense. Diverse examples are presented, showing that the deformed model is smooth and the presented canonical model is easy to use. Yong-Jin Liu 0001, Matthew M. F. Yuen |
CAD/Graphics | 1 |
| 2007 | Developable Strip Approximation of Parametric Surfaces with Global Error BoundsabstractDevelopable surfaces have many desired properties in manufacturing process. Since most existing CAD systems utilize parametric surfaces as the design primitive, there is a great demand in industry to convert a parametric surface within a prescribed global error bound into developable patches. In this work we propose a simple and efficient solution to approximate a general parametric surface with a minimum set of C0-joint developable strips. The key contribution of the proposed algorithm is that, several global optimization problems are elegantly solved in a sequence that offers a controllable global error bound on the developable surface approximation. Experimental results are presented to demonstrate the effectiveness and stability of the proposed algorithm. Yong-Jin Liu 0001, Yukun Lai, Shi-Min Hu 0001 |
PG | 1 |
| 2007 | Modeling dynamic developable meshes by the Hamilton principle
Yong-Jin Liu 0001, Kai Tang 0001, Ajay Joneja |
Comput. Aided Des. | 1 |
| 2007 | Handling degenerate cases in exact geodesic computation on triangle meshes
Yong-Jin Liu 0001, Qian-Yi Zhou, Shi-Min Hu 0001 |
Vis. Comput. | 1 |
| 2006 | Dynamic Medial Axes of Planar Shapes
Kai Tang 0001, Yong-Jin Liu 0001 |
Computer Graphics International | 2 |
| 2006 | An Efficient Implementation of RBF-Based Progressive Point-Sampled Geometry
Yong-Jin Liu 0001, Kai Tang 0001, Ajay Joneja |
GMP | 1 |
| 2005 | Sketch-based free-form shape modelling with a fast and stable numerical engine
Yong-Jin Liu 0001, Kai Tang 0001, Ajay Joneja |
Comput. Graph. | 1 |
| 2005 | An optimization algorithm for free-form surface partitioning based on weighted gaussian image
Kai Tang 0001, Yong-Jin Liu 0001 |
Graph. Model. | 2 |
| 2004 | Efficient and Stable Numerical Algorithms on Equilibrium Equations for Geometric ModelingabstractIn this paper the applications of equilibrium equation to geometric modeling is exploited and efficient numerical algorithms are proposed for solving the equilibrium equation. First we show that from diverse geometric modeling applications the equilibrium system can be extracted as the central framework. Second, by exploiting in-depth the special structures inherent in the geometric applications, we present simplified analytic solutions to the resulting geometric equilibrium equations via system decomposition. Finally, given the observation that the geometric equilibrium systems are extremely sensitive to both perturbations in input data and round off errors, efficient, stable and accurate numerical algorithms are proposed. Yong-Jin Liu 0001, Kai Tang 0001, Matthew M. F. Yuen |
GMP | 1 |
| 2004 | Multiresolution Free Form Object Modeling with Point Sampled Geometry
Yong-Jin Liu 0001, Kai Tang 0001, Matthew M. F. Yuen |
J. Comput. Sci. Technol. | 1 |
| 2004 | A geometric method for determining intersection relations between a movable convex object and a set of planar polygonsabstractIn this paper, we investigate how to topologically and geometrically characterize the intersection relations between a movable convex polygon A and a set /spl Xi/ of possibly overlapping polygons fixed in the plane. More specifically, a subset /spl Phi//spl sube//spl Xi/ is called an intersection relation if there exists a placement of A that intersects, and only intersects, /spl Phi/. The objective of this paper is to design an efficient algorithm that finds a finite and discrete representation of all of the intersection relations between A and /spl Xi/. Past related research only focuses on the complexity of the free space of the configuration space between A and /spl Xi/ and how to move or place an object in this free space. However, there are many applications that require the knowledge of not only the free space, but also the intersection relations. Examples are presented to demonstrate the rich applications of the formulated problem on intersection relations. Kai Tang 0001, Yong-Jin Liu 0001 |
IEEE Trans. Robotics | 2 |
| 2003 | Maximal intersection of spherical polygons by an arc with applications to 4-axis machining
Kai Tang 0001, Yong-Jin Liu 0001 |
Comput. Aided Des. | 2 |
| 2003 | Optimized triangle mesh reconstruction from unstructured points
Yong-Jin Liu 0001, Matthew M. F. Yuen |
Vis. Comput. | 1 |
| 2003 | Manifold-guaranteed out-of-core simplification of large meshes with controlled topological type
Yong-Jin Liu 0001, Matthew M. F. Yuen, Kai Tang 0001 |
Vis. Comput. | 1 |
| 2002 | A feature-based approach for individualized human head modeling
Yong-Jin Liu 0001, Matthew M. F. Yuen, Shan Xiong |
Vis. Comput. | 1 |