EDBT 2026 Demo / reviewers in the wild / expert
Minjing Yu
dblp:173/0901
· DBLP profile ↗
24ranked-venue papers
8as first author
16since 2021 · last 2025
0000-0001-6755-2027ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VAction: A Lightweight and Integrated VR Training System for Authentic Film-Shooting Experience
Che Qu, Minjing Yu, Chao Zhou 0012, Yuntao Wang 0001, Yu-Hui Wen, Yuanchun Shi, Yong-Jin Liu 0001 |
CHI | 3 |
| 2025 | Exploring the Influence of Profile Picture Styles on Empathy and Identity Recognition in Social MediaabstractEmpathy and identity recognition are two core social interaction factors that greatly affect efficiency and effectiveness. The selection of a profile picture in social media is crucial as it serves as a visual representation of virtual identity. It not only reflects the user's personality but also influences the level of empathy and identification that others feel toward them. However, the potential impact of profile pictures with different styles (e.g., real faces, cartoon faces, and landscape images) on empathy and identity recognition is still unclear. To explore its effects, a controlled laboratory experiment and an ecological online experiment were conducted. Participants were shown a picture each time, informed to imagine interacting with the person using it as his/her profile picture, and instructed to rate an item from the basic empathy scale (BES) based on it. After rating all pictures, users then completed an identity recognition task. Results show that participants’ empathy scores for users with cartoon or real face profile pictures are greater than those with landscape profile pictures. In addition, participants performed better in identity recognition for users with real face or landscape images as profile pictures than for those with cartoon face profile pictures. Moreover, users of social media often make social categorizations (i.e., in-group/out-group categorization) based on the social identities expressed by their profile pictures. Our results also indicate that the affective empathy scale rating score was positively associated with the degree to which users of the corresponding profile pictures were categorized as in-group members. Minjing Yu, Xinge Liu, Chao Zhou 0012, Xinxin Du, Jenny Sheng, Yong-Jin Liu 0001 |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2025 | PCKRF: Point Cloud Completion and Keypoint Refinement With Fusion Data for 6D Pose EstimationabstractSome robust point cloud registration approaches with controllable pose refinement magnitude, such as ICP and its variants, are commonly used to improve 6D pose estimation accuracy. However, the effectiveness of these methods gradually diminishes with the advancement of deep learning techniques and the enhancement of initial pose accuracy, primarily due to their lack of specific design for pose refinement. In this paper, we propose Point Cloud Completion and Keypoint Refinement with Fusion Data (PCKRF), a new pose refinement pipeline for 6D pose estimation. The pipeline consists of two steps. First, it completes the input point clouds via a novel pose-sensitive point completion network. The network uses both local and global features with pose information during point completion. Then, it registers the completed object point cloud with the corresponding target point cloud by our proposed Color supported Iterative KeyPoint (CIKP) method. The CIKP method introduces color information into registration and registers a point cloud around each keypoint to increase stability. The PCKRF pipeline can be integrated with existing popular 6D pose estimation methods, such as the full flow bidirectional fusion network, to further improve their pose estimation accuracy. Experiments demonstrate that our method exhibits superior stability compared to existing approaches when optimizing initial poses with relatively high precision. Notably, the results indicate that our method effectively complements most existing pose estimation techniques, leading to improved performance in most cases. Furthermore, our method achieves promising results even in challenging scenarios involving textureless and symmetrical objects. Yiheng Han, Irvin Haozhe Zhan, Long Zeng 0001, Yu-Ping Wang 0001, Ran Yi 0002, Minjing Yu, Matthieu Lin, Jenny Sheng, Yong-Jin Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | ECAvatar: 3D Avatar Facial Animation with Controllable Identity and Emotion
Minjing Yu, Delong Pang, Ziwen Kang, Zhiyao Sun, Tian Lv, Jenny Sheng, Ran Yi 0002, Yu-Hui Wen, Yong-Jin Liu 0001 |
ACM Multimedia | 1 |
| 2024 | VisHanfu: An Interactive System for the Promotion of Hanfu Knowledge via Cross-Shaped Flat Structure
Minjing Yu, Lingzhi Zeng, Xinxin Du, Jenny Sheng, Qiantian Liao, Yong-Jin Liu 0001 |
ACM Multimedia | 1 |
| 2024 | MakeBronze: An interactive system to promote Chinese bronze culture in children through hands-on experience with lost-wax casting
Minjing Yu, Li Wang 0131, Mingxu Cai, Mengrui Zhang, Chun Yu, Xing-Dong Yang, Jiawan Zhang |
Int. J. Hum. Comput. Stud. | 1 |
| 2024 | MFDAN: Multi-Level Flow-Driven Attention Network for Micro-Expression RecognitionabstractFacial expressions are an essential part of human emotional communication, and micro-expressions (MEs), as transient and imperceptible non-verbal signals, can potentially reveal real human emotions. However, subtle motion variations, limited and unbalanced samples make micro-expression recognition (MER) challenging. In this paper, we design a novel dual-branch learning framework of multi-level flow-driven attention for micro-expression recognition (MFDAN), which innovatively integrates optical flow prior to guide the attention learning in the image encoding branch, enabling the model to focus on the most discriminative facial regions for subtle motion patterns. Firstly, we extract optical flow information by an optical flow encoding module. Then, in the image coding module, we construct a Transformer structure containing an optical flow-driven attention mechanism, which can effectively locate the interest region of micro-expressions in the image according to the position information of optical flow to capture more sensitive and fine-grained micro-expressions. By interoperating prior knowledge with data learning, and introducing the Dropkey operation and Focal Loss, our method can handle subtle micro-expression features on small imbalanced datasets. Through extensive experiments on three independent datasets and a composite database, including SMIC-HS, SAMM, and CASME II, robust leave-one-subject-out (LOSO) evaluation results show that our method outperforms state-of-the-art methods especially on the composite database. Junli Zhao, Ran Yi 0002, Minjing Yu, Fuqing Duan, Zhenkuan Pan 0001, Yong-Jin Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | DiffPoseTalk: Speech-Driven Stylistic 3D Facial Animation and Head Pose Generation via Diffusion ModelsabstractThe generation of stylistic 3D facial animations driven by speech presents a significant challenge as it requires learning a many-to-many mapping between speech, style, and the corresponding natural facial motion. However, existing methods either employ a deterministic model for speech-to-motion mapping or encode the style using a one-hot encoding scheme. Notably, the one-hot encoding approach fails to capture the complexity of the style and thus limits generalization ability. In this paper, we propose DiffPoseTalk, a generative framework based on the diffusion model combined with a style encoder that extracts style embeddings from short reference videos. During inference, we employ classifier-free guidance to guide the generation process based on the speech and style. In particular, our style includes the generation of head poses, thereby enhancing user perception. Additionally, we address the shortage of scanned 3D talking face data by training our model on reconstructed 3DMM parameters from a high-quality, in-the-wild audio-visual dataset. Extensive experiments and user study demonstrate that our approach outperforms state-of-the-art methods. The code and dataset are at https://diffposetalk.github.io. Zhiyao Sun, Tian Lv, Matthieu Lin, Jenny Sheng, Yu-Hui Wen, Minjing Yu, Yong-Jin Liu 0001 |
ACM Trans. Graph. | 7 |
| 2023 | An Easy-to-Build Modular Robot Implementation of Chain-Based Physical Transformation for STEM Education
Minjing Yu, Jeffrey Too Chuan Tan, Yong-Jin Liu 0001 |
CAD/Graphics | 1 |
| 2023 | 4D facial analysis: A survey of datasets, algorithms and applications
Yong-Jin Liu 0001, Baodong Wang, Lin Gao 0004, Junli Zhao, Ran Yi 0002, Minjing Yu, Zhenkuan Pan 0001, Xianfeng Gu |
Comput. Graph. | 6 |
| 2023 | Generation of virtual digital human for customer service industry
Yanan Sun 0006, Zhiyao Sun, Yu-Hui Wen, Tian Lv, Minjing Yu, Ran Yi 0002, Lin Gao 0004, Yong-Jin Liu 0001 |
Comput. Graph. | 6 |
| 2023 | SparseDGCNN: Recognizing Emotion From Multichannel EEG SignalsabstractEmotion recognition from EEG signals has attracted much attention in affective computing. Recently, a novel dynamic graph convolutional neural network (DGCNN) model was proposed, which simultaneously optimized the network parameters and a weighted graph$G$characterizing the strength of functional relation between each pair of two electrodes in the EEG recording equipment. In this article, we propose a sparse DGCNN model which modifies DGCNN by imposing a sparseness constraint on$G$and improves the emotion recognition performance. Our work is based on an important observation: the tomography study reveals that different brain regions sampled by EEG electrodes may be related to different functions of the brain and then the functional relations among electrodes are possibly highly localized and sparse. However, introducing sparseness constraint into the graph$G$makes the loss function of sparse DGCNN non-differentiable at some singular points. To ensure that the training process of sparse DGCNN converges, we apply the forward-backward splitting method. To evaluate the performance of sparse DGCNN, we compare it with four representative recognition methods (SVM, DBN, GELM and DGCNN). In addition to comparing different recognition methods, our experiments also compare different features and spectral bands, including EEG features in time-frequency domain (DE, PSD, DASM, RASM, ASM and DCAU on different bands) extracted from four representative EEG datasets (SEED, DEAP, DREAMER, and CMEED). The results show that (1) sparse DGCNN has consistently better accuracy than representative methods and has a good scalability, and (2) DE, PSD, and ASM features on$\gamma$band convey most discriminative emotional information, and fusion of separate features and frequency bands can improve recognition performance. Minjing Yu, Yong-Jin Liu 0001, Guozhen Zhao, Dan Zhang 0014, Wenming Zheng |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | 3D-CariGAN: An End-to-End Solution to 3D Caricature Generation From Normal Face PhotosabstractCaricature is a type of artistic style of human faces that attracts considerable attention in the entertainment industry. So far a few 3D caricature generation methods exist and all of them require some caricature information (e.g., a caricature sketch or 2D caricature) as input. This kind of input, however, is difficult to provide by non-professional users. In this paper, we propose an end-to-end deep neural network model that generates high-quality 3D caricatures directly from a normal 2D face photo. The most challenging issue for our system is that the source domain of face photos (characterized by normal 2D faces) is significantly different from the target domain of 3D caricatures (characterized by 3D exaggerated face shapes and textures). To address this challenge, we: (1) build a large dataset of 5,343 3D caricature meshes and use it to establish a PCA model in the 3D caricature shape space; (2) reconstruct a normal full 3D head from the input face photo and use its PCA representation in the 3D caricature shape space to establish correspondences between the input photo and 3D caricature shape; and (3) propose a novel character loss and a novel caricature loss based on previous psychological studies on caricatures. Experiments including a novel two-level user study show that our system can generate high-quality 3D caricatures directly from normal face photos. Zipeng Ye, Mengfei Xia, Yanan Sun 0006, Ran Yi 0002, Minjing Yu, Juyong Zhang, Yukun Lai, Yong-Jin Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | We Can Do More to Save Guqin: Design and Evaluate Interactive Systems to Make Guqin More Accessible to the General PublicabstractGuqin is a plucked seven-string traditional Chinese musical instrument that exists for over 3,000 years. However, as an Intangible World Cultural Heritage, the inheritance of Guqin and its culture in modern society is in deep danger. According to our study with 1,006 Chinese worldwide, Guqin as an instrument is not well-known and barely accessible. To better promote Guqin, we developed two interactive systems: VirGuqin and MRGuqin. VirGuqin was developed using a low-cost motion tracking device and was tested in a museum. 89% of 308 participants expressed an increase in interest in learning Guqin after using our system. MRGuqin was developed as a mixed reality learning environment to reduce the entry barrier to Guqin, and was tested by 16 participants, allowing them to learn Guqin significantly faster and perform better than the current practice. Our study demonstrates how technology can be used to help the inheritance of this dying art. Minjing Yu, Chun Yu, Xiaoguang Ma, Xing-Dong Yang, Jiawan Zhang |
CHI | 1 |
| 2021 | Feature-Aware Uniform Tessellations on Video Manifold for Content-Sensitive SupervoxelsabstractOver-segmenting a video into supervoxels has strong potential to reduce the complexity of downstream computer vision applications. Content-sensitive supervoxels (CSSs) are typically smaller in content-dense regions (i.e., with high variation of appearance and/or motion) and larger in content-sparse regions. In this paper, we propose to compute feature-aware CSSs (FCSSs) that are regularly shaped 3D primitive volumes well aligned with local object/region/motion boundaries in video. To compute FCSSs, we map a video to a 3D manifold embedded in a combined color and spatiotemporal space, in which the volume elements of video manifold give a good measure of the video content density. Then any uniform tessellation on video manifold can induce CSS in the video. Our idea is that among all possible uniform tessellations on the video manifold, FCSS finds one whose cell boundaries well align with local video boundaries. To achieve this goal, we propose a novel restricted centroidal Voronoi tessellation method that simultaneously minimizes the tessellation energy (leading to uniform cells in the tessellation) and maximizes the average boundary distance (leading to good local feature alignment). Theoretically our method has an optimal competitive ratio O(1), and its time and space complexities are O(NK) and O(N+K) for computing K supervoxels in an N-voxel video. We also present a simple extension of FCSS to streaming FCSS for processing long videos that cannot be loaded into main memory at once. We evaluate FCSS, streaming FCSS and ten representative supervoxel methods on four video datasets and two novel video applications. The results show that our method simultaneously achieves state-of-the-art performance with respect to various evaluation criteria. Ran Yi 0002, Zipeng Ye, Wang Zhao 0001, Minjing Yu, Yukun Lai, Yong-Jin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Automatic Sitting Pose Generation for Ergonomic Ratings of ChairsabstractHuman poses play a critical role in human-centric product design. Despite considerable researches on pose synthesis and pose-driven product design, most of them adopt the simple stick figure model that captures only skeletons rather than real body geometries and do not link human poses to the environment (e.g., chairs for sitting). This paper focuses on user-tailored ergonomic design and rating of chairs using scanned human geometries. Fully utilizing the anthropometric information of the human models, our method considers more ergonomic guidelines of chair design (such as pressure distribution and support intensity) and links the geometry of 3D chair models and human-to-chair interactions into the pose deformation constraints of the human avatars. The core of our method is a pose generation algorithm which rigs the user's successive poses through coarse- and fine-level pose deformations. We define a non-linear energy function with contact, collision, and joint limit terms, and solve it using a hill-climbing algorithm. The fitting results allow us to quantitatively evaluate the chair model in terms of various ergonomic criteria. Our method is flexible and effective and can be applied to users with varying body shapes and a wide range of chairs. Moreover, the proposed technique can be easily extended to other furniture, such as desk, bed, and cabinet. Extensive evaluations and a user study demonstrate the efficiency and advantages of the proposed virtual fitting method. Given that our method avoids tedious on-site trying, facilitates the exploration/evaluation of various chair products, and provides valuable feedback for the designers and manufacturers to deliver customized products, it is ideal for online shopping of chairs. Aihua Mao, Zhenfeng Xie, Minjing Yu, Yong-Jin Liu 0001, Ying He 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2019 | Fast Computation of Content-Sensitive Superpixels and Supervoxels Using Q-DistancesabstractState-of-the-art researches model the data of images and videos as low-dimensional manifolds and generate superpixels/supervoxels in a content-sensitive way, which is achieved by computing geodesic centroidal Voronoi tessellation (GCVT) on manifolds. However, computing exact GCVTs is slow due to computationally expensive geodesic distances. In this paper, we propose a much faster queue-based graph distance (called q-distance). Our key idea is that for manifold regions in which q-distances are different from geodesic distances, GCVT is prone to placing more generators in them, and therefore after few iterations, the q-distance-induced tessellation is an exact GCVT. This idea works well in practice and we also prove it theoretically under moderate assumption. Our method is simple and easy to implement. It runs 6-8 times faster than state-of-the-art GCVT computation, and has an optimal approximation ratio O(1) and a linear time complexity O(N) for N-pixel images or N-voxel videos. A thorough evaluation of 31 superpixel methods on five image datasets and 8 supervoxel methods on four video datasets shows that our method consistently achieves the best over-segmentation accuracy. We also demonstrate the advantage of our method on one image and two video applications. Zipeng Ye, Ran Yi 0002, Minjing Yu, Yong-Jin Liu 0001, Ying He 0001 |
ICCV | 3 |
| 2019 | LineUp: Computing Chain-Based Physical TransformationabstractIn this article, we introduce a novel method that can generate a sequence of physical transformations between 3D models with different shape and topology. Feasible transformations are realized on a chain structure with connected components that are 3D printed. Collision-free motions are computed to transform between different configurations of the 3D printed chain structure. To realize the transformation between different 3D models, we first voxelize these input models into a similar number of voxels. The challenging part of our approach is to generate a simple path—as a chain configuration to connect most voxels. A layer-based algorithm is developed with theoretical guarantee of the existence and the path length. We find that collision-free motion sequence can always be generated when using a straight line as the intermediate configuration of transformation. The effectiveness of our method is demonstrated by both the simulation and the experimental tests taken on 3D printed chains. Minjing Yu, Zipeng Ye, Yong-Jin Liu 0001, Ying He 0001, Charlie C. L. Wang |
ACM Trans. Graph. | 1 |
| 2018 | Intrinsic Manifold SLIC: A Simple and Efficient Method for Computing Content-Sensitive SuperpixelsabstractSuperpixels are perceptually meaningful atomic regions that can effectively capture image features. Among various methods for computing uniform superpixels, simple linear iterative clustering (SLIC) is popular due to its simplicity and high performance. In this paper, we extend SLIC to compute content-sensitive superpixels, i.e., small superpixels in content-dense regions with high intensity or colour variation and large superpixels in content-sparse regions. Rather than using the conventional SLIC method that clusters pixels in , we map the input image to a 2-dimensional manifold , whose area elements are a good measure of the content density in . We propose a simple method, called intrinsic manifold SLIC (IMSLIC), for computing a geodesic centroidal Voronoi tessellation (GCVT)-a uniform tessellation-on , which induces the content-sensitive superpixels in . In contrast to the existing algorithms, IMSLIC characterizes the content sensitivity by measuring areas of Voronoi cells on . Using a simple and fast approximation to a closed-form solution, the method can compute the GCVT at a very low cost and guarantees that all Voronoi cells are simply connected. We thoroughly evaluate IMSLIC and compare it with eleven representative methods on the BSDS500 dataset and seven representative methods on the NYUV2 dataset. Computational results show that IMSLIC outperforms existing methods in terms of commonly used quality measures pertaining to superpixels such as compactness, adherence to boundaries, and achievable segmentation accuracy. We also evaluate IMSLIC and seven representative methods in an image contour closure application, and the results on two datasets, WHD and WSD, show that IMSLIC achieves the best foreground segmentation performance. Yong-Jin Liu 0001, Minjing Yu, Bing-Jun Li, Ying He 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Real-Time Movie-Induced Discrete Emotion Recognition from EEG SignalsabstractRecognition of a human's continuous emotional states in real time plays an important role in machine emotional intelligence and human-machine interaction. Existing real-time emotion recognition systems use stimuli with low ecological validity (e.g., picture, sound) to elicit emotions and to recognise only valence and arousal. To overcome these limitations, in this paper, we construct a standardised database of 16 emotional film clips that were selected from over one thousand film excerpts. Based on emotional categories that are induced by these film clips, we propose a real-time movie-induced emotion recognition system for identifying an individual's emotional states through the analysis of brain waves. Thirty participants took part in this study and watched 16 standardised film clips that characterise real-life emotional experiences and target seven discrete emotions and neutrality. Our system uses a 2-s window and a 50 percent overlap between two consecutive windows to segment the EEG signals. Emotional states, including not only the valence and arousal dimensions but also similar discrete emotions in the valence-arousal coordinate space, are predicted in each window. Our real-time system achieves an overall accuracy of 92.26 percent in recognising high-arousal and valenced emotions from neutrality and 86.63 percent in recognising positive from negative emotions. Moreover, our system classifies three positive emotions (joy, amusement, tenderness) with an average of 86.43 percent accuracy and four negative emotions (anger, disgust, fear, sadness) with an average of 65.09 percent accuracy. These results demonstrate the advantage over the existing state-of-the-art real-time emotion recognition systems from EEG signals in terms of classification accuracy and the ability to recognise similar discrete emotions that are close in the valence-arousal coordinate space. Yong-Jin Liu 0001, Minjing Yu, Guozhen Zhao, Jinjing Song, Yan Ge 0007, Yuanchun Shi |
IEEE Trans. Affect. Comput. | 2 |
| 2016 | Manifold SLIC: A Fast Method to Compute Content-Sensitive SuperpixelsabstractSuperpixels are perceptually meaningful atomic regions that can effectively capture image features. Among various methods for computing uniform superpixels, simple linear iterative clustering (SLIC) is popular due to its simplicity and high performance. In this paper, we extend SLIC to compute content-sensitive superpixels, i.e., small superpixels in content-dense regions (e.g., with high intensity or color variation) and large superpixels in content-sparse regions. Rather than the conventional SLIC method that clusters pixels in ℝ5, we map the image I to a 2-dimensional manifold M ⊂ ℝ5, whose area elements are a good measure of the content density in I. We propose an efficient method to compute restricted centroidal Voronoi tessellation (RCVT) - a uniform tessellation - on M, which induces the content-sensitive superpixels in I. Unlike other algorithms that characterize content-sensitivity by geodesic distances, manifold SLIC tackles the problem by measuring areas of Voronoi cells on M, which can be computed at a very low cost. As a result, it runs 10 times faster than the state-of-the-art content-sensitive superpixels algorithm. We evaluate manifold SLIC and seven representative methods on the BSDS500 benchmark and observe that our method outperforms the existing methods. Yong-Jin Liu 0001, Cheng-Chi Yu, Minjing Yu, Ying He 0001 |
CVPR | 3 |
| 2016 | Cognitive mechanism related to line drawings and its applications in intelligent process of visual media: a survey
Yong-Jin Liu 0001, Minjing Yu, Qiu-Fang Fu, Ye Liu 0010, Lexing Xie |
Frontiers Comput. Sci. | 2 |
| 2016 | A PMJ-inspired cognitive framework for natural scene categorization in line drawings
Minjing Yu, Yong-Jin Liu 0001, Qiu-Fang Fu, Xiaolan Fu |
Neurocomputing | 1 |
| 2016 | A Robust Divide and Conquer Algorithm for Progressive Medial Axes of Planar ShapesabstractThe medial axis is an important shape representation that finds a wide range of applications in shape analysis. For large-scale shapes of high resolution, a progressive medial axis representation that starts with the lowest resolution and gradually adds more details is desired. In this paper, we propose a fast and robust geometric algorithm that computes progressive medial axes of a large-scale planar shape. The key ingredient of our method is a novel structural analysis of merging medial axes of two planar shapes along a shared boundary. Our method is robust by separating the analysis of topological structure from numerical computation. Our method is also fast and we show that the time complexity of merging two medial axes is$O(n\;\log n_v)$, where$n$is the number of total boundary generators,$n_v$is strictly smaller than$n$and behaves as a small constant in all our experiments. Experiments on large-scale polygonal data and comparison with state-of-the-art methods show the efficiency and effectiveness of the proposed method. Yong-Jin Liu 0001, Cheng-Chi Yu, Minjing Yu, Kai Tang 0001, Deok-Soo Kim |
IEEE Trans. Vis. Comput. Graph. | 3 |