EDBT 2026 Demo / reviewers in the wild / expert
Jiaqing Liu
dblp:226/2675
· DBLP profile ↗
29ranked-venue papers
3as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GenesisFunc: Multi-Agent Data Generation for Accurate and Generalizable Function-CallingabstractHao-Xiang Xu, Chong Deng, Jiaqing Liu, Wen Wang, Qian Chen, Lujia Bao, Xiangang Li, Zhen-Hua Ling. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hao-Xiang Xu, Chong Deng, Jiaqing Liu, Wen Wang 0001, Qian Chen 0003, Lujia Bao, Xiangang Li, Zhen-Hua Ling |
ACL (1) | 3 |
| 2026 | Circuit2Yarn: From Planar Circuits to Electronic Yarns for Textile-Based InteractionsabstractSmart yarns hold the potential to transform everyday textiles into functional platforms, yet current methods remain constrained. These include conductive yarns, made from silver or stainless steel, which retain the feel of conventional yarns but offer limited functions, and PCB-based solutions, which add capability at the cost of bulk and rigidity. We present Circuit2Yarn, a fabrication framework that transforms planar printed circuits into flexible yarns by rolling copper-traced TPU films with soldered surface-mount components, preserving the capabilities of rigid electronics while producing yarn-like forms suitable for textile integration. We demonstrate yarns as small as 0.8 mm that integrate LEDs and sensors, including temperature, humidity, light, IMU, and capacitive sensing modules, enabling applications ranging from smart garments and interactive musical instruments to responsive tea bags. Characterization confirms durability under bending/stretching. By rolling planar circuits into yarns, Circuit2Yarn paves the way toward comfortable, multifunctional, and interactive textiles in everyday life. Zhechen Zhao, Tianhong Catherine Yu, Jiaqing Liu, Huaishu Peng, Yiyue Luo, Zhihan Zhang 0002, Tingyu Cheng |
CHI | 4 |
| 2026 | RA3-FDA: Resource-adaptive federated domain adaptation with dual heterogeneity awareness for EEG-based depression detection
Siyang Song, Huaning Wang, Jiewei Jiang, Dongmei Jiang, Jie Zhang 0028, Prayag Tiwari, Jiaqing Liu |
Expert Syst. Appl. | 13 |
| 2026 | Enhancing Depression Detection Using Pretrained Multi-modal Sentiment Analysis Models with Deep Prefix TuningabstractDepression, a pervasive mental health condition, affects millions globally, challenging early and accurate diagnosis due to its subtle and varied manifestations. Recognizing the critical link between emotional dysregulation and depressive symptoms, our research introduces a pioneering training paradigm that integrates sentiment analysis with depression detection. This approach is motivated by the potential of sentiment data to enrich models with a deeper understanding of emotional states, crucial for identifying depressive patterns. To leverage the nuanced sentiment information without compromising the pretrained model’s integrity, we employ deep prefix tuning. This novel technique allows for targeted model refinement, ensuring that the valuable pretrained structures are not overshadowed by the sparse and specific nature of depression-related data. The empirical results demonstrate superior performance across standard benchmarks, setting a new precedent for multimodal depression detection. Shiyu Teng, Jiaqing Liu, Shurong Chai, Hao Sun 0013, Tomoko Tateyama, Lanfen Lin, Yen-Wei Chen 0001 |
ACM Trans. Comput. Heal. | 2 |
| 2026 | OPCR: Continual generalized category discovery via orthogonal prototypes and confidence-aware label refinement
Ningge Hu, Xiao Li 0008, Jiaqing Liu, Xuezheng Fan |
Neurocomputing | 5 |
| 2026 | One framework to rule them all: Unifying multimodal tasks with LLM neural-tuning
Hao Sun 0013, Yu Song 0008, Jiaqing Liu, Jihong Hu, Yen-Wei Chen 0001, Lanfen Lin |
Pattern Recognit. | 3 |
| 2026 | Multi-binary network with feature disentanglement for Open-set Specific Emitter Identification
Ningge Hu, Jiaqing Liu |
Signal Process. | 5 |
| 2026 | Disentangled Multimodal Tuning and Interaction for Human Perception UnderstandingabstractUnderstanding human perceptions poses a significant multimodal challenge for computers, involving textual, acoustic, and visual signals. Recently, large language models (LLMs) have garnered great attention, leading to numerous methods aimed at efficiently fine-tuning pretrained models for multimodal downstream tasks. However, there remains a scarcity of techniques that prioritize modality-invariant and -specific information during parameter-efficient tuning, despite evidence from previous studies showcasing the effectiveness of modality disentangling. To address this gap, we propose a novel multimodal tuning approach for LLMs, termed Disentangled Multimodal Tuning and Interaction. Specifically, we evaluate the independence among different modalities and disentangle corresponding modality-invariant and specific components, which are subsequently leveraged for prompt tuning. Following tuning, a newly designed independence-guided cross-attention module is introduced for modality interaction, where the attention mechanism is decoupled and bolstered with independence from the modality-disentangling process. This approach not only enables LLMs to efficiently assimilate information from various modalities but also cultivates an awareness of both modality-invariant and specific information. Compared to previous methods, our approach facilitates modality interaction at a more granular level, resulting in enhanced performance. We validate our method through experiments on four public datasets, demonstrating significant performance improvements. Hao Sun 0013, Ziwei Niu, Jiaqing Liu, Yen-Wei Chen 0001, Lanfen Lin |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Recording for Eyes, Not Echoing to Ears: Contextualized Spoken-to-Written Conversion of ASR TranscriptsabstractAutomatic Speech Recognition (ASR) transcripts exhibit recognition errors and various spoken language phenomena such as disfluencies, ungrammatical sentences, and incomplete sentences, hence suffering from poor readability. To improve readability, we propose a Contextualized Spoken-to-Written conversion (CoS2W) task to address ASR and grammar errors and also transfer the informal text into the formal style with content preserved, utilizing contexts and auxiliary information. This task naturally matches the in-context learning capabilities of Large Language Models (LLMs). To facilitate comprehensive comparisons of various LLMs, we construct a document-level Spoken-to-Written conversion of ASR Transcripts Benchmark (SWAB) dataset. Using SWAB, we study the impact of different granularity levels on the CoS2W performance, and propose methods to exploit contexts and auxiliary information to enhance the outputs. Experimental results reveal that LLMs have the potential to excel in the CoS2W task, particularly in grammaticality and formality, our methods achieve effective understanding of contexts and auxiliary information by LLMs. We further investigate the effectiveness of using LLMs as evaluators and find that LLM evaluators show strong correlations with human evaluations on rankings of faithfulness and formality, which validates the reliability of LLM evaluators for the CoS2W task. Jiaqing Liu, Chong Deng, Shilin Zhou 0002, Qian Chen 0003, Wen Wang 0001 |
AAAI | 1 |
| 2025 | OmniFlatten: An End-to-end GPT Model for Seamless Voice ConversationabstractQinglin Zhang, Luyao Cheng, Chong Deng, Qian Chen, Wen Wang, Siqi Zheng, Jiaqing Liu, Hai Yu, Chao-Hong Tan, Zhihao Du, ShiLiang Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Luyao Cheng, Chong Deng, Qian Chen 0003, Wen Wang 0001, Jiaqing Liu, Chao-Hong Tan, Zhihao Du, Shiliang Zhang |
ACL (1) | 7 |
| 2025 | ECG Necklace: Low-power Wireless Necklace for Continuous ECG monitoring
Qiuyue Xue, Eric Steven Martin, Jiaqing Liu, Ruiqing Wang, Antonio Glenn, Richard Li 0002, Vikram Iyer, Shwetak N. Patel |
CHI | 3 |
| 2025 | Enhanced Multimodal Depression Detection With Emotion PromptsabstractDepression is a pervasive mental health disorder that remains frequently undiagnosed and untreated due to societal barriers and the subjective nature of its symptoms. Leveraging recent advances in large language models (LLMs), we propose a novel depression detection pipeline that generates emotion prompts tailored to individual data, enhancing detection accuracy. Our approach integrates cross-modality fusion via cross attention mechanisms to combine depressive and emotional features, creating a comprehensive representation of depression indicators. Evaluated on the E-DAIC and EATD datasets, our method outperforms state-of-the-art techniques, demonstrating its potential for more precise emotion-based depression detection. Shiyu Teng, Jiaqing Liu, Hao Sun 0013, Shurong Chai, Tomoko Tateyama, Lanfen Lin, Yen-Wei Chen 0001 |
ICASSP | 2 |
| 2025 | Clinical Data-Driven Retrieval-Augmented Model for Lung Nodule Malignancy Prediction
Ruibo Hou, Shurong Chai, Rahul Kumar Jain 0001, Yinhao Li 0002, Jiaqing Liu, Shiyu Teng, Lanfen Lin, Yen-Wei Chen 0001 |
MICCAI (10) | 5 |
| 2025 | Multimodal Sentiment Analysis With Mutual Information-Based Disentangled Representation LearningabstractMultimodal sentiment analysis seeks to utilize various types of signals to identify underlying emotions and sentiments. A key challenge in this field lies in multimodal representation learning, which aims to develop effective methods for integrating multimodal features into cohesive representations. Recent advancements include two notable approaches: one focuses on decomposing multimodal features into modality-invariant and -specific components, while the other emphasizes the use of mutual information to enhance the fusion of modalities. Both strategies have demonstrated effectiveness and yielded remarkable results. In this paper, we propose a novel learning framework that combines the strengths of these two approaches, termed mutual information-based disentangled multimodal representation learning. Our approach involves estimating different types of information during feature extraction and fusion stages. Specifically, we quantitatively assess and adjust the proportions of modality-invariant, -specific, and -complementary information during feature extraction. Subsequently, during fusion, we evaluate the amount of information retained by each modality in the fused representation. We employ mutual information or conditional mutual information to estimate each type of information content. By reconciling the proportions of these different types of information, our approach achieves state-of-the-art performance on popular sentiment analysis benchmarks, including CMU-MOSI and CMU-MOSEI. Hao Sun 0013, Ziwei Niu, Hongyi Wang 0002, Xinyao Yu 0003, Jiaqing Liu, Yen-Wei Chen 0001, Lanfen Lin |
IEEE Trans. Affect. Comput. | 5 |
| 2024 | Loss Masking Is Not Needed In Decoder-Only Transformer For Discrete-Token-Based ASRabstractRecently, unified speech-text models, such as SpeechGPT, VioLA, and AudioPaLM, have achieved remarkable performance on various speech tasks. These models discretize speech signals into tokens (speech discretization) and use a shared vocabulary for both text and speech tokens. Then they train a single decoder-only Transformer on a mixture of speech tasks. However, these models rely on the Loss Masking strategy for the ASR task, which ignores the dependency among speech tokens. In this paper, we propose to model speech tokens in an autoregressive way, similar to text. We find that applying the conventional cross-entropy loss on input speech tokens does not consistently improve the ASR performance over the Loss Masking approach. To address this issue, we propose a novel approach denoted Smoothed Label Distillation (SLD), which applies a KL divergence loss with smoothed labels on speech tokens. Our experiments show that SLD effectively models speech tokens and outperforms Loss Masking for decoder-only Transformers in ASR tasks with different speech discretization methods1. Qian Chen 0003, Wen Wang 0001, Shiliang Zhang, Chong Deng, Jiaqing Liu, Chong Zhang 0003 |
ICASSP | 9 |
| 2024 | Ladder Fine-tuning Approach for SAM Integrating Complementary NetworkabstractRecently, foundation models have been introduced demonstrating various tasks in the field of computer vision. These models such as Segment Anything Model (SAM) are generalized models trained using huge datasets. Currently, ongoing research focuses on exploring the effective utilization of these generalized models for Specific domains, such as medical imaging. However, in medical imaging, the lack of training samples due to privacy concerns and other factors presents a major challenge for applying these generalized models to medical image segmentation task. To address this issue, the effective fine tuning of these models is crucial to ensure their optimal utilization. In this study, we propose to combine a complementary Convolutional Neural Network (CNN) along with the standard SAM network for medical image segmentation. To reduce the burden of fine tuning large foundation model and implement cost-efficient training scheme, we focus only on fine-tuning the additional CNN network and SAM decoder part. This strategy significantly reduces training time and achieves competitive results on publicly available dataset. The code is available at ">https://github.com/11yxk/SAM-LST . Shurong Chai, Rahul Kumar Jain 0001, Shiyu Teng, Jiaqing Liu, Yinhao Li 0002, Tomoko Tateyama, Yen-Wei Chen 0001 |
KES | 4 |
| 2024 | A Novel Adaptive Hypergraph Neural Network for Enhancing Medical Image Segmentation
Shurong Chai, Rahul Kumar Jain 0001, Shaocong Mo, Jiaqing Liu, Yinhao Li 0002, Tomoko Tateyama, Lanfen Lin, Yen-Wei Chen 0001 |
MICCAI (9) | 4 |
| 2024 | A motion-aware and temporal-enhanced Spatial-Temporal Graph Convolutional Network for skeleton-based human action segmentation
Shurong Chai, Rahul Kumar Jain 0001, Jiaqing Liu, Shiyu Teng, Tomoko Tateyama, Yinhao Li 0002, Yen-Wei Chen 0001 |
Neurocomputing | 3 |
| 2023 | Ditto: A Simple and Efficient Approach to Improve Sentence EmbeddingsabstractPrior studies diagnose the anisotropy problem in sentence representations from pre-trained language models, e.g., BERT, without finetuning.Our analysis reveals that the sentence embeddings from BERT suffer from a bias towards uninformative words, limiting the performance in semantic textual similarity (STS) tasks.To address this bias, we propose a simple and efficient unsupervised approach, Diagonal Attention Pooling (Ditto), which weights words with model-based importance estimations and computes the weighted average of word representations from pre-trained models as sentence embeddings.Ditto can be easily applied to any pre-trained language model as a postprocessing operation.Compared to prior sentence embedding approaches, Ditto does not add parameters nor requires any learning.Empirical evaluations demonstrate that our proposed Ditto can alleviate the anisotropy problem and improve various pre-trained models on the STS benchmarks. 1 Qian Chen 0003, Wen Wang 0001, Chong Deng, Jiaqing Liu, Chong Zhang 0003 |
EMNLP | 7 |
| 2023 | Improving Long Document Topic Segmentation Models With Enhanced Coherence ModelingabstractTopic segmentation is critical for obtaining structured documents and improving downstream tasks such as information retrieval.Due to its ability of automatically exploring clues of topic shift from abundant labeled data, recent supervised neural models have greatly promoted the development of long document topic segmentation, but leaving the deeper relationship between coherence and topic segmentation underexplored.Therefore, this paper enhances the ability of supervised models to capture coherence from both logical structure and semantic similarity perspectives to further improve the topic segmentation performance, proposing Topic-aware Sentence Structure Prediction (TSSP) and Contrastive Semantic Similarity Learning (CSSL).Specifically, the TSSP task is proposed to force the model to comprehend structural information by learning the original relations between adjacent sentences in a disarrayed document, which is constructed by jointly disrupting the original document at topic and sentence levels.Moreover, we utilize inter-and intra-topic information to construct contrastive samples and design the CSSL objective to ensure that the sentences representations in the same topic have higher similarity, while those in different topics are less similar.Extensive experiments show that Longformer with our approach significantly outperforms state-of-the-art (SOTA) methods.Our approach improves F 1 of SOTA by 3.42 (73.74 → 77.16) and improves P k by 1.11 points (15.0 → 13.89) on WIKI-727K and achieves an average relative reduction of 4.3% on P k on WikiSection.The average relative P k drop of 8.38% on two out-of-domain datasets also demonstrates the robustness of our approach 1 . Chong Deng, Jiaqing Liu, Qian Chen 0003, Wen Wang 0001 |
EMNLP | 4 |
| 2023 | Meeting Action Item Detection with Regularized Context ModelingabstractMeetings are increasingly important for collaborations. Action items in meeting transcripts are crucial for managing post-meeting to-do tasks, which usually are summarized laboriously. The Action Item Detection task aims to automatically detect meeting content associated with action items. However, datasets manually annotated with action item detection labels are scarce and in small scale. We construct and release the first Chinese meeting corpus with manual action item annotations1. In addition, we propose a Context-Drop approach to utilize both local and global contexts by contrastive learning, and achieve better accuracy and robustness for action item detection. We also propose a Lightweight Model Ensemble method to exploit different pre-trained models2. Experimental results on our Chinese meeting corpus and the English AMI corpus demonstrate the effectiveness of the proposed approaches. Jiaqing Liu, Chong Deng, Qian Chen 0003, Wen Wang 0001 |
ICASSP | 1 |
| 2023 | Overview of the ICASSP 2023 General Meeting Understanding and Generation Challenge (MUG)abstractICASSP2023 General Meeting Understanding and Generation Challenge (MUG) focuses on prompting a wide range of spoken language processing (SLP) research on meeting transcripts, as SLP applications are critical to improve users’ efficiency in grasping important information in meetings. MUG includes five tracks, including topic segmentation, topic-level and session-level extractive summarization, topic title generation, keyphrase extraction, and action item detection. To facilitate MUG, we construct and release a large-scale meeting dataset, the AliMeeting4MUG Corpus. We review the dataset, track settings and baselines, and summarize the challenge results and major techniques used in the submissions. Chong Deng, Jiaqing Liu, Qian Chen 0003, Wen Wang 0001, Zhijie Yan, Jinglin Liu, Yi Ren 0006, Zhou Zhao 0001 |
ICASSP | 3 |
| 2023 | MUG: A General Meeting Understanding and Generation BenchmarkabstractListening to long video/audio recordings from video conferencing and online courses for acquiring information is extremely inefficient. Even after ASR systems transcribe recordings into long-form spoken language documents, reading ASR transcripts only partly speeds up seeking information. It has been observed that a range of NLP applications, such as keyphrase extraction, topic segmentation, and summarization, significantly improve users’ efficiency in grasping important information. The meeting scenario is among the most valuable scenarios for deploying these spoken language processing (SLP) capabilities. However, the lack of large-scale public meeting datasets annotated for these SLP tasks severely hinders their advancement. To prompt SLP advancement, we establish a large-scale general Meeting Understanding and Generation Benchmark (MUG) to benchmark the performance of a wide range of SLP tasks, including topic segmentation, topic-level and session-level extractive summarization and topic title generation, keyphrase extraction, and action item detection. To facilitate the MUG benchmark, we construct and release a large-scale meeting dataset for comprehensive long-form SLP development, the AliMeeting4MUG Corpus, which consists of 654 recorded Mandarin meeting sessions with diverse topic coverage, with manual annotations for SLP tasks on manual transcripts of meeting recordings. To the best of our knowledge, the AliMeeting4MUG Corpus is so far the largest meeting corpus in scale and facilitates most SLP tasks. In this paper, we provide a detailed introduction of this corpus, SLP tasks and evaluation methods, baseline systems and their performance1. Chong Deng, Jiaqing Liu, Qian Chen 0003, Wen Wang 0001, Zhijie Yan, Jinglin Liu, Yi Ren 0006, Zhou Zhao 0001 |
ICASSP | 3 |
| 2023 | Intelligent UAS-Edge-Server Collaboration and Orchestration in Disaster Response ManagementabstractUnmanned aerial systems (UAS) consist of a swarm of unmanned aerial vehicles (UAVs) with edge resources and collaboration with ground-control-servers (GCS) are useful for heavy computation use cases e.g., traffic management, public safety, and disaster response management. Inefficient setups and collaboration decisions, often stemming from edge/cloud network misconfigurations, can lead to suboptimal resource utilization and delayed response times. In this paper, we present a novel scheme for (soft) real-time learning-based UAS-Edge-Server collaboration and orchestration strategies to achieve pertinent allocations of both computation resources and communication strategies. Our approach includes i) policy-based pre-application collaboration and benchmark analysis as well as ii) learning-based multi-agent deep Q-network (DQN) algorithm that optimizes UAV swarm trajectories during application. Evaluation results demonstrate that our policy-based approach Pareto-optimally trade-off performance (e.g., accuracy, streaming) and disaster response time. In addition, our DQN approach significantly enhances edge-cloud resource cooperation, improving network performance metrics like throughput and round-trip time by a minimum of 12% compared to state-of-the-art edge-internet-of-things (EIoT) collaboration algorithms. Furthermore, through real-world emulations, we illustrate how our orchestration attains 87% of the Oracle baseline network throughput performance while maintaining a comparable disaster response time for various video analytics-based disaster scenarios. Chengyi Qu, Chaise Ballotti, Daniel De Sousa, Jiaqing Liu |
WETICE | 4 |
| 2022 | CubeMLP: An MLP-based Model for Multimodal Sentiment Analysis and Depression EstimationabstractMultimodal sentiment analysis and depression estimation are two important research topics that aim to predict human mental states using multimodal data. Previous research has focused on developing effective fusion strategies for exchanging and integrating mind-related information from different modalities. Some MLP-based techniques have recently achieved considerable success in a variety of computer vision tasks. Inspired by this, we explore multimodal approaches with a feature-mixing perspective in this study. To this end, we introduce CubeMLP, a multimodal feature processing framework based entirely on MLP. CubeMLP consists of three independent MLP units, each of which has two affine transformations. CubeMLP accepts all relevant modality features as input and mixes them across three axes. After extracting the characteristics using CubeMLP, the mixed multimodal features are flattened for task predictions. Our experiments are conducted on sentiment analysis datasets: CMU-MOSI and CMU-MOSEI, and depression estimation dataset: AVEC2019. The results show that CubeMLP can achieve state-of-the-art performance with a much lower computing cost. Hao Sun 0013, Hongyi Wang 0002, Jiaqing Liu, Yen-Wei Chen 0001, Lanfen Lin |
ACM Multimedia | 3 |
| 2022 | A multi-head pseudo nodes based spatial-temporal graph convolutional network for emotion perception from GAIT
Shurong Chai, Jiaqing Liu, Rahul Kumar Jain 0001, Tomoko Tateyama, Yutaro Iwamoto, Lanfen Lin, Yen-Wei Chen 0001 |
Neurocomputing | 2 |
| 2021 | Sequence Model with Self-Adaptive Sliding Window for Efficient Spoken Document SegmentationabstractTranscripts generated by automatic speech recognition (ASR) systems for spoken documents lack structural annotations such as paragraphs, significantly reducing their readability. Automatically predicting paragraph segmentation for spoken documents may both improve readability and downstream NLP performance such as summarization and machine reading comprehension. We propose a sequence model with self-adaptive sliding window for accurate and efficient paragraph segmentation. We also propose an approach to exploit pho-netic information, which significantly improves robustness of spoken document segmentation to ASR errors. Evaluations are conducted on the English Wiki-727K document seg-mentation benchmark, a Chinese Wikipedia-based document segmentation dataset we created, and an in-house Chinese spoken document dataset. Our proposed model outperforms the state-of-the-art (SOTA) model based on the same BERT-Base, increasing segmentation F1 on the English benchmark by 4.2 points and on Chinese datasets by 4.3-10.1 points, while reducing inference time to less than 1/6 of inference time of the current SOTA. Qian Chen 0003, Yali Li 0001, Jiaqing Liu, Wen Wang 0001 |
ASRU | 4 |
| 2019 | An Improved Hand Gesture Recognition with Two-Stage Convolution Neural Networks Using a Hand Color Image and its Pseudo-Depth ImageabstractRobust hand gesture recognition has been playing a significant role in the field of human-computer interaction for a long time, but it is still full of challenges due to many accept such as cluttered backgrounds and hand self-occlusion. With the help of depth information, depth-based methods have better performance, but the depth cameras are not as widely used and affordable as color cameras. Therefore, in this paper, we propose a two-stage deep convolutional neural network (CNN) architecture for accurate color-based hand gesture recognition. The first stage performs generation of pseudo-depth hand images from color images and the second stage recognizes hand gesture classes using both the color image and its pseudo-depth hand image. The generation stage architecture is based on an image-to-image translation network. In the recognition stage, a two-stream CNN architecture with color image and its pseudo depth image is proposed to improve the color image-based recognition performance. We also propose two strategies in two-stream fusion: feature fusion and committee fusion. To validate our approach, we construct a new dataset called MaHG-RGBD dataset. Experiments demonstrate that our approach significantly improves the performance in RGB-only recognition for hand gestures. Jiaqing Liu, Kotaro Furusawa, Tomoko Tateyama, Yutaro Iwamoto, Yen-Wei Chen 0001 |
ICIP | 1 |
| 2018 | Dense Optical Flow Variation Based 3D Face Reconstruction from Monocular VideoabstractThis paper presents a method for reconstructing 3D face expressions from monocular video sequences. Unlike previous approaches we don't require any prior face models, nor a large collection of images with diverse variation of poses and illuminations. Instead, we leverage a monocular video sequence without any restrictions. We formulate the 3D face reconstruction as an energy minimization problem integrated with dense optical flow variation, as rigid as possible(ARAP) constraint, spatial and temporal constraints. This paper offers the first dense optical flow variational approach to the problem of 3D reconstruction of non-rigid face expressions from a monocular video. Dense optical flow variation cost substitutes for photo consistency cost to enhance the reconstruction of exaggerated expressions. A generic 3D face template mesh and a simple 3D warping algorithm allow us to reconstruct a true 3D face mesh, relax the constraints of diverse views or illuminations and also avoid the dependency of the quality of prior face models, such as the facial expressions or face races varieties limitation. Finally, we use a per-pixel shape-from-shading(SFS) algorithm to estimate the fine-scale geometry details such as wrinkles to further improve the reconstruction fidelity. Given unconstrained monocular RGB videos, our method reconstructs wrinkle-level 3D face model, without the need for any prior models or diverse capture conditions. Shan Wang 0012, Xukun Shen, Jiaqing Liu |
ICIP | 3 |