VLDB 2026 Research / reviewers in the wild / expert
Junfeng Yao
dblp:00/2981
· DBLP profile ↗
126ranked-venue papers
9as first author
97since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 63 · 6 first-author · 44 since 2021Artificial intelligence and machine learning · 56 · 1 first-author · 47 since 2021Human-computer interaction and ubiquitous computing · 14 · 6 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 6 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MGD: Mesh-guided Gaussians with Diffusion Priors for Dynamic Objects Reconstruction from Monocular RGB-D VideoabstractReconstructing dynamic objects from monocular RGB-D video is critical for advancing 3D vision applications and enhancing user experience. However, monocular RGB-D video provides limited 3D observations, making the reconstruction of unobserved regions highly under-constrained. Despite recent advances that combine neural implicit surfaces with diffusion models, the inherent limitations of implicit representations and the lack of effective guidance in diffusion priors lead to blurry appearance and inaccurate geometry in dynamic object reconstruction. To address the issue, we present MGD, which leverages scene-adaptive diffusion priors and Mesh-guided Gaussians for realistic rendering and geometrically accurate reconstruction of dynamic objects, including unobserved regions. The dynamic 3D objects reconstructed by MGD are represented using our proposed Mesh-guided Gaussians, which leverage global and local Gaussians to capture large-scale deformations and fine-grained appearance details, respectively. Additionally, in order to utilize depth information, we integrate a depth ControlNet into the diffusion model and conduct scene-adaptive fine-tuning. We design a self-generated image-pair strategy to produce image pairs used for fine-tuning. Extensive experiments demonstrate that MGD achieves state-of-the-art performance in both high-fidelity reconstruction and structural completeness, while maintaining real-time efficiency during training and rendering. Weixing Xie, Jintian Li, Bingchuan Li, Yanchen Lin, Junfeng Yao |
AAAI | 7 |
| 2026 | Benchmarking and Learning Real-World Customer Service DialogueabstractTianhong Gao, Jundong Shen, Jiapeng Wang, Bei Shi, Ying Ju, Junfeng Yao, Huiyu Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tianhong Gao, Jundong Shen, Bei Shi, Junfeng Yao, Huiyu Yu |
ACL (1) | 6 |
| 2026 | Cultural relic image restoration using two-stage transformer-CNN framework
Xing Wu 0001, Deyu Gao, Junfeng Yao, Quan Qian |
Appl. Intell. | 4 |
| 2026 | HIB-MSRL: Rethinking multi-scale representation learning: A hierarchical information bottleneck perspective
Junfeng Yao, Chunyan Li 0002 |
Expert Syst. Appl. | 4 |
| 2026 | SPC: Self-supervised point cloud completion
Xing Wu 0001, Junfeng Yao, Chenhao Shang, Quan Qian |
Neural Networks | 3 |
| 2025 | SimRP: Syntactic and Semantic Similarity Retrieval Prompting Enhances Aspect Sentiment Quad PredictionabstractAspect Sentiment Quad Prediction (ASQP) is the most complex subtask of Aspect-based Sentiment Analysis (ABSA), aiming to predict all sentiment quadruples within the given sentence. Due to the complexity of sentence syntaxes and the diversity of sentiment expressions, generative methods gradually become the mainstream approach in ASQP. However, existing generative models are constrained in the effectiveness of demonstrations. Semantically similar demonstrations help in judging sentiment categories and polarities but may confuse the model in recognizing aspect and opinion terms, which are more related to sentence syntaxes. To this end, we first develop Syn2Vec, a method for calculating syntactic vectors to support the retrieval of syntactically similar demonstrations. Then, we propose Syntactic and Semantic Similarity Retrieval Prompting (SimRP) to construct effective prompts by retrieving the most related demonstrations that are syntactically and semantically similar. With these related demonstrations, pre-trained generative models, especially Large Language Models (LLMs), can fully release their potential to recognize sentiment quadruples. Extensive experiments in Supervised Fine-Tuning (SFT) and In-context Learning (ICL) paradigms demonstrate the effectiveness of SimRP. Furthermore, we find that LLMs' capabilities in ASQP are severely underestimated by biased data annotations and the exact matching metric. We propose a novel constituent subtree-based fuzzy metric for more accurate and rational quadruple recognition. Zhongquan Jian, Yanhao Chen 0002, Jiajian Li, Shaopan Wang, Xiangjian Zeng, Junfeng Yao, Xinying An, Qingqiang Wu 0001 |
AAAI | 6 |
| 2025 | Locate-and-Focus: Enhancing Terminology Translation in Speech Language ModelsabstractDirect speech translation (ST) has garnered increasing attention nowadays, yet the accurate translation of terminology within utterances remains a great challenge. In this regard, current studies mainly concentrate on leveraging various translation knowledge into ST models. However, these methods often struggle with interference from irrelevant noise and can not fully utilize the translation knowledge. To address these issues, in this paper, we propose a novel Locate-and-Focus method for terminology translation. It first effectively locates the speech clips containing terminologies within the utterance to construct translation knowledge, minimizing irrelevant information for the ST model. Subsequently, it associates the translation knowledge with the utterance and hypothesis from both audio and textual modalities, allowing the ST model to better focus on translation knowledge during translation. Experimental results across various datasets demonstrate that our method effectively locates terminologies within utterances and enhances the success rate of terminology translation, while maintaining robust general translation performance. Suhang Wu, Jialong Tang, Pei Zhang 0011, Baosong Yang, Junhui Li 0001, Junfeng Yao, Min Zhang 0005, Jinsong Su |
ACL (1) | 7 |
| 2025 | AGCL: Aspect Graph Construction and Learning for Aspect-level Sentiment ClassificationabstractPrior studies on Aspect-level Sentiment Classification (ALSC) emphasize modeling interrelationships among aspects and contexts but overlook the crucial role of aspects themselves as essential domain knowledge. To this end, we propose AGCL, a novel Aspect Graph Construction and Learning method, aimed at furnishing the model with finely tuned aspect information to bolster its task-understanding ability. AGCL’s pivotal innovations reside in Aspect Graph Construction (AGC) and Aspect Graph Learning (AGL), where AGC harnesses intrinsic aspect connections to construct the domain aspect graph, and then AGL iteratively updates the introduced aspect graph to enhance its domain expertise, making it more suitable for the ALSC task. Hence, this domain aspect graph can serve as a bridge connecting unseen aspects with seen aspects, thereby enhancing the model’s generalization capability. Experiment results on three widely used datasets demonstrate the significance of aspect information for ALSC and highlight AGL’s superiority in aspect learning, surpassing state-of-the-art baselines greatly. Code is available at https://github.com/jian-projects/agcl. Zhongquan Jian, Daihang Wu, Shaopan Wang, Junfeng Yao, Meihong Wang, Qingqiang Wu 0001 |
COLING | 5 |
| 2025 | CMU-Flownet: Exploring Point Cloud Scene Flow Estimation in Occluded Scenario
Jingze Chen, Zerui Tang, Qiqin Lin, Junfeng Yao |
CVM (2) | 5 |
| 2025 | Learnable Infinite Taylor Gaussian for Dynamic View RenderingabstractCapturing the temporal evolution of Gaussian properties such as position, rotation, and scale is a challenging task due to the vast number of time-varying parameters and the limited photometric data available, which generally results in convergence issues, making it difficult to find an optimal solution. While feeding all inputs into an end-to- end neural network can effectively model complex temporal dynamics, this approach lacks explicit supervision and struggles to generate high-quality transformation fields. On the other hand, using time-conditioned polynomial functions to model Gaussian trajectories and orientations provides a more explicit and interpretable solution, but requires significant handcrafted effort and lacks generalizability across diverse scenes. To overcome these limitations, this paper introduces a novel approach based on a learnable infinite Taylor Formula to model the temporal evolution of Gaussians. This method offers both the flexibility of an implicit network-based approach and the interpretability of explicit polynomial functions, allowing for more robust and generalizable modeling of Gaussian dynamics across various dynamic scenes. Extensive experiments on dynamic novel view rendering tasks are conducted on public datasets, demonstrating that the proposed method achieves state-of-the-art performance in this domain. More information is available on our project page (https://ellisonking.github.io/TaylorGaussian). Haoye Dong, Junfeng Yao, Gim Hee Lee |
CVPR | 6 |
| 2025 | Emotional Knowledge Self-Distillation in DialogueabstractRecognizing emotions in dialogues is vital for effective human-computer interaction, yet remains a challenging task in Natural Language Processing (NLP). Previous studies in Emotion Recognition in Conversation (ERC) have primarily focused on contextual features, while overlooking the importance of emotional features in emotion recognition. To address this gap, we focus on the role of emotional features in ERC and propose a novel method, Emotional Knowledge Self-Distillation (EmoKSD1), to enhance the model’s emotional sensitivity. In EmoKSD, utterances are enriched with implicit ⟨mask⟩ tokens to represent conveyed emotions, allowing the distillation of emotional knowledge from explicit emotional tokens to implicit ⟨mask⟩ tokens, thereby enhancing the model’s ability to perceive subtle emotions within the dialogue. Through thorough evaluations on two public ERC datasets (i.e., IEMOCAP and MELD) using proposed coarse-grained utterance distillation and fine-grained token distillation techniques, EmoKSD demonstrates superior performance compared to existing methods, highlighting the significance of emotional features in ERC. Zhongquan Jian, Weichao Wu, Junfeng Yao, Meihong Wang, Qingqiang Wu 0001 |
ICASSP | 4 |
| 2025 | Curriculum Contrastive Learning for Aspect-based Sentiment AnalysisabstractPre-trained Language Models (PLMs) have achieved remarkable performance in various Natural Language Processing (NLP) tasks, including Aspect-based Sentiment Analysis (ABSA). Therefore, numerous ABSA models based on PLMs have been proposed, primarily focusing on module design to exploit the inherent connections between aspects and contexts. However, the core factor driving performance improvements, the PLM’s powerful semantic understanding capabilities, has not been fully considered, raising the question of how to further unlock their potential for downstream tasks. To this end, we introduce a novel training strategy, called CCL1, which integrates the strengths of Curriculum Learning (CurL) and Contrastive Learning (ConL) to facilitate the learning of robust feature representations. For the ABSA task, we use aspect similarities to develop the CurL strategy, grouping samples with similar aspects into batches. This allows ConL to learn more robust representations by providing related samples within each batch. The superiority of CCL is demonstrated through extensive experiments on two public ABSA datasets, with ablation studies validating the effectiveness of combining CurL and ConL in enhancing aspect understanding. Zhongquan Jian, Daihang Wu, Xiangjian Zeng, Junfeng Yao, Meihong Wang, Qingqiang Wu 0001 |
ICASSP | 4 |
| 2025 | PD-SDF: Dynamic Surface Reconstruction Based on Plane Decomposition for Single View RGB-D VideosabstractSurface reconstruction of dynamic scenes from single view videos is a challenging task due to the highly ill-posed and under-constrained nature. Existing single view reconstruction methods suffer from severe quality issues, such as surface distortion and mesh adherison. In this paper, we propose an efficient dynamic representation network, PD-SDF, which consists of a 4D motion field and a 3D geometry field. The explicit disentanglement of motion and geometry based on planar factorization guarantees the mesh consistency across frames. Specifically, to address the mesh adhesion problem, we design a depth-guided sampling strategy to focus on optimizing the SDF field near the object surface. Due to insufficient geometric cues, we design various regularization strategies to constrain smoothness and topological correcness of the scene geometry. Extensive experiments show that our method outperforms existing methods in both appearance and geometry reconstruction. The project page: https://pd-sdf.github.io/ Weixing Xie, Junfeng Yao, Shaoqi Wu, Youhong Peng, Mengyuan Ge |
ICASSP | 3 |
| 2025 | Enhancing Information Extraction with METORIE: A Metaphor and Trap-Based Dataset for Cross-Domain Fine-TuningabstractThis research proposes the METORIE dataset1, a novel resource designed to improve the reasoning capabilities of large language models (LLMs), such as LLaMA3 and GLM4, in information extraction (IE) tasks. The METORIE dataset is derived from brain teasers that incorporate complex logical and metaphorical elements and is designed to train LLMs to navigate intricate reasoning paths and interpret layered expressions. Our findings demonstrate that the METORIE dataset markedly enhances LLMs’ performance across both general and specialized IE tasks. The results of fine-tuning with the METORIE dataset, mixed with a small number of IE datasets, are close to, if not exceeding, those of LLMs of the same parametric size on IE tasks using much larger datasets. Through controlled experiments, we establish that metaphors of medium complexity optimize IE performance, while higher complexities tend to overstretch LLMs’ inference limits. METORIE-fine-tuned LLMs also demonstrate exceptional performance in legal and medical domains, suggesting that enhanced metaphor understanding and logical deduction are key to improving LLMs’ adaptability and efficiency in vertical domains. Zhengyuan Pan, Yilian Peng, Zhongquan Jian, Yanhao Chen 0002, Wentao Qiu, Haonan Ma, Junfeng Yao, Meihong Wang, Qingqiang Wu 0001 |
ICASSP | 7 |
| 2025 | LDG: Lightweight Deformable 3D Gaussians for Single-View Dynamic Scene ReconstructionabstractRecent deformable 3D Gaussians methods achieve high-quality reconstruction and real-time rendering. However, they require multi-view information and are not applicable to single-view dynamic scenes captured from mobile phones. Additionally, the high-dimensional hidden layer of deformation MLP and the excessive number of Gaussian primitives and attributes impose significant storage pressure, greatly limiting their practical application. To address the issues, we propose a novel Lightweight Deformable 3D Gaussians teacher-student framework. Specifically, we initialize Gaussian primitives with an initialization strategy designed for single-view scenes, and then optimize the teacher model using color and depth information. For the trained teacher model, we distill deformation MLP, prune Gaussian primitives and Gaussian attributes, and finally obtain a student model with low storage and high efficiency. Public benchmark experiments demonstrate the effectiveness of our framework, showing a compression rate exceeding 4× while maintaining satisfactory rendering quality. Project page: https://poyoki.github.io/ldg/. Youhong Peng, Weixing Xie, Shaoqi Wu, Junfeng Yao |
ICASSP | 7 |
| 2025 | LP-Gaussians: Learnable Parametric Gaussian Splatting for Efficient Dynamic Reconstruction of Single-View ScenesabstractWith the popularity of short video platforms, the number of single-view videos has increased significantly. Existing NeRF-based methods can reconstruct dynamic scenes in a single-view setting, but slow rendering speed and low rendering quality limit their practical applications. To address these challenges, we propose a fast single-view scenes reconstruction framework based on 3D Gaussian Splatting. Our method uses point clouds obtained with depth priors as the Gaussian initialization and introduces learnable parametric functions to model the time-dependent deformation of Gaussians. The explicit deformation modeling for Gaussians significantly reduces training and rendering time. Furthermore, to improve the rendering quality of challenging areas, we adopt an adaptive sampling strategy to densify Gaussians. For occlusion problems from single-view videos, we design a smooth loss function to restore the color of the occluded areas. Experimental results demonstrate that our method significantly reduces training time, enhances rendering quality, and accelerates rendering speed. Project page: https://github.com/LPGaussians. Shaoqi Wu, Weixing Xie, Youhong Peng, Jiawei Yao, Junfeng Yao |
ICASSP | 6 |
| 2025 | Supervised Exploratory Learning for Long-Tailed Visual Recognition
Zhongquan Jian, Yanhao Chen 0002, Junfeng Yao, Meihong Wang, Qingqiang Wu 0001 |
ICCV | 4 |
| 2025 | Integrating Multi-modal Interaction with AR Garment Design for Educational Applications
Hengheng Zhao, Chaohui Yang, Junfeng Yao, Ruolong Wang |
ICXR | 3 |
| 2025 | The First MPDD Challenge: Multimodal Personality-aware Depression DetectionabstractDepression is a widespread mental health issue affecting diverse age groups, with notable prevalence among college students and the elderly. However, existing datasets and detection methods primarily focus on young adults, neglecting the broader age spectrum and individual differences that influence depression manifestation. Current approaches often establish a direct mapping between multimodal data and depression indicators, failing to capture the complexity and diversity of depression across individuals. This challenge includes two tracks based on age-specific subsets: Track 1 uses the MPDD-Elderly dataset for detecting depression in older adults, and Track 2 uses the MPDD-Young dataset for detecting depression in younger participants. The Multimodal Personality-aware Depression Detection (MPDD) Challenge aims to address this gap by incorporating multimodal data alongside individual difference factors. We provide a baseline model that fuses audio and video modalities with individual difference information to detect depression manifestations in diverse populations. This challenge aims to promote the development of more personalized and accurate de pression detection methods, advancing mental health research and fostering inclusive detection systems. More details are available on the official challenge website: https://hacilab.github.io/MPDDChallenge.github.io. Changzeng Fu, Zelin Fu, Qi Zhang 0124, Xinhe Kuang, Jiacheng Dong, Kaifeng Su, Yikai Su, Junfeng Yao, Yuliang Zhao, Shiqi Zhao 0001, Siyang Song, Yuichiro Yoshikawa, Björn W. Schuller, Hiroshi Ishiguro |
ACM Multimedia | 9 |
| 2025 | Sera: Separated Coarse-to-fine Representation Alignment for Cross-subject EEG-based Emotion RecognitionabstractNeuropsychology-inspired models have been utilized in recent advances in EEG emotion recognition, such as convolutional networks for spatial features and Transformers for temporal dependencies. While these methods benefit from domain knowledge like frequency-band features and spatial correlations, most overlook the fundamental fact that EEG signals are complex mixtures of neural source activities recorded at the scalp. EEG signals presenting challenges for emotion recognition, particularly in cross-subject scenarios due to significant inter-subject variance. Inspired by neurophysiological principles, we propose a novel framework, named Sera, for EEG-based emotion recognition that explicitly separates source activities and aligns representations across subjects. Sera introduces two key components: (1) a variational autoencoder (VAE) with multiple multi-stage decoders (M2VAE) designed to disentangle EEG signals into independent sources, mimicking the neural generation process, and (2) a coarse-to-fine representation alignment block (CFRA) to mitigate subject-to-subject variability. The coarse alignment employs adversarial training with a domain discriminator, while the fine-grained alignment matches covariance matrices to capture temporal correlations within EEG segments. Extensive experiments demonstrate that Sera outperforms the state-of-the-art methods with improvements ranging from 1% to 5%, averaging 3.14% and 3.05% on the DEAP and DREAMER datasets, respectively, confirming its effectiveness and neurophysiological grounding. The code is available at: https://github.com/JZH98/Sera-code. Meiyan Xu, Ziyu Jia, Yong Li 0032, Xinliang Zhou, Junfeng Yao, Yi Ding 0012 |
ACM Multimedia | 8 |
| 2025 | Gloss Matters: Unlocking the Potential of Non-Autoregressive Sign Language TranslationabstractWhile non-autoregressive sign language translation (NASLT) has the advantage in inference speed, the translation quality of NASLT models lags significantly behind that of the state-of-the-art (SOTA) autoregressive sign language translation (ASLT) models. To bridge the quality gap, we exploit glosses to unlock the potential of NASLT models. Concretely, we propose Gloss-enhanced Levenshtein Transformer (GLevT) for sign language translation (SLT), which takes glosses as initial sequences for editing into texts. In particular, to alleviate the inconsistency between training and inference of GLevT, which is introduced by glosses, we propose a dual-centric learning policy and a keyframe-based gloss replacement method for training, further improving the translation quality of GLevT. Experiments on CSL-Daily demonstrate that GLevT outperforms other NASLT models by approximately 4 points in BLEU and ROUGE scores, while achieving performance comparable to the SOTA ASLT models with a 3.46~5.26× inference speed-up. Furthermore, we extend GLevT to gloss-free SLT, achieving performance comparable to SOTA large models, despite having only 49M parameters. We release code at https://github.com/XMUDeepLIT/GLevT. Zhiwei He 0002, Kangjie Zheng, Liangying Shao, Junfeng Yao, Jinsong Su |
ACM Multimedia | 6 |
| 2025 | PianoPal: A Robotic Multimedia System for Interactive Piano Instruction Based on Q-Learning and Real-Time Feedback
Junfeng Yao |
MMM (3) | 2 |
| 2025 | Optimization of Single-Track Train Schedules with Cyclic Operation StrategiesabstractThis paper proposes an integrated mixed-integer programming model, termed the Single-Track Railway Cyclic Scheduling Model (SRCSM), for constructing optimized periodic timetables (operating on a recurring 24-hour cycle) for bidirectional single-track railway systems. The SRCSM enhances operational efficiency by precisely considering train arrival/departure times, safety headways, platform track allocations, and meet/pass operations. Validation using real-world data demonstrates its practical applicability, flexibility, and scalability for dynamic timetable optimization, offering a robust tool for improving operational efficiency and safety in single-track railway operations. Xing Wu 0001, Deyu Gao, Junfeng Yao, Quan Qian |
SoMeT | 3 |
| 2025 | NPGCL: neighbor enhancement and embedding perturbation with graph contrastive learning for recommendation
Xing Wu 0001, Junfeng Yao, Quan Qian |
Appl. Intell. | 3 |
| 2025 | Scnet: spectral convolutional networks for multivariate time series classificationabstractAbstract With the widespread application of time series data, the study of classification techniques has become an important topic. Although existing multivariate time series classification (MTSC) methods have made progress, they often rely on one-dimensional (1D) time series, which limits their ability to capture complex temporal dynamics and multiscale features. To address these challenges, a Spectral Convolutional Network (SCNet) is introduced in this work. SCNet effectively transforms 1D time series data into the frequency domain using an enhanced Discrete Fourier Transform (enhanced_DFT), revealing periodicity and key frequency components while reshaping the data into a two-dimensional (2D) time series for better representation. Furthermore, it uses a Spectral Energy Prioritization method to optimize frequency domain energy distribution and a multiscale convolutional module to capture features at different scales, improving the model’s ability to analyze short-term and long-term trends. To validate the effectiveness and superiority, we conducted extensive experiments on 10 sub-datasets from the well-known UEA dataset. The results show that our proposed SCNet achieved the highest average accuracy of 74.3%, which is 2.2% higher than the current state-of-the-art models, demonstrating its potential for practical application and efficiency in MTSC task. Xing Wu 0001, Junfeng Yao, Quan Qian |
Appl. Intell. | 3 |
| 2025 | SFNS: Spatial-frequency image noise suppression for low-power industrial cone-beam computed tomography
Xing Wu 0001, Junfeng Yao, Quan Qian, Shouwei Gao |
Appl. Intell. | 3 |
| 2025 | RBF-MAT: Computing medial axis transform from point clouds by optimizing radial basis functions
Mengyuan Ge, Junfeng Yao, Baorong Yang, Ningna Wang, Zhonggui Chen, Xiaohu Guo |
Comput. Aided Geom. Des. | 2 |
| 2025 | Mitigating the negative impact of over-association for conversational query production
Ante Wang, Linfeng Song, Zijun Min, Xiaoli Wang 0002, Junfeng Yao, Jinsong Su |
Inf. Process. Manag. | 6 |
| 2025 | Multi-views Emotional Knowledge Extraction for Emotion Recognition in Conversation
Zhongquan Jian, Daihang Wu, Shaopan Wang, Jiezhou He, Junfeng Yao, Kunhong Liu 0001, Qingqiang Wu 0001 |
Knowl. Based Syst. | 5 |
| 2025 | Endo-HDR: Dynamic endoscopic reconstruction with deformable 3D Gaussians and hierarchical depth regularization
Weixing Xie, Qingqi Hong, Junfeng Yao, Shaoqi Wu, Rongzhou Zhou, Xiaohu Guo |
Knowl. Based Syst. | 4 |
| 2025 | OVST: online video stabilization with two-stage training transformer
Xing Wu 0001, Junfeng Yao, Quan Qian, Yike Guo |
Neural Comput. Appl. | 5 |
| 2025 | Aspect sentiment learning for Aspect-Level Sentiment Classification
Zhongquan Jian, Jiajian Li, Meihong Wang, Junfeng Yao, Qingqiang Wu 0001 |
Neural Networks | 4 |
| 2025 | Towards better text image machine translation with multimodal codebook and multi-stage training
Zhibin Lan, Junfeng Yao, Degen Huang, Jinsong Su |
Neural Networks | 4 |
| 2025 | TDAG: A multi-agent framework based on dynamic Task Decomposition and Agent Generation
Yaoxiang Wang, Zhiyong Wu 0003, Junfeng Yao, Jinsong Su |
Neural Networks | 3 |
| 2024 | Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized RehearsalabstractJianheng Huang, Leyang Cui, Ante Wang, Chengyi Yang, Xinting Liao, Linfeng Song, Junfeng Yao, Jinsong Su. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Jianheng Huang, Leyang Cui, Ante Wang, Xinting Liao, Linfeng Song, Junfeng Yao, Jinsong Su |
ACL (1) | 7 |
| 2024 | S2CCT: Self-Supervised Collaborative CNN-Transformer for Few-shot Medical Image SegmentationabstractSelf-supervised pre-training followed by fine-tuning is a potent paradigm for few-shot learning, leveraging extensive unlabeled data with remarkable efficacy. Current self-supervised methods often lean towards Vision Transformers (ViTs) rather than CNN-Transformer hybrid architectures, which generally demonstrate superior performance. However, this reliance on ViTs can lead to poor perception of local features by the model. The challenge lies in designing a suitable proxy task for hybrid architectures like CNN-Transformers, which have significant structural differences. Additionally, the current organization of CNN-Transformer hybrid backbones is often sequential, hindering collaboration during pre-training and the acquisition of robust representations. To address these issues, we propose Self-Supervised Collaborative CNN-Transformer (S2CCT) for few-shot medical image segmentation. This framework introduces three innovative designs: (1) a composite proxy task based on image masking and image super-resolution tailored for CNN-Transformer hybrid architectures, enabling the backbone to acquire robust representations during pre-training that can be transferred to downstream tasks; (2) a parallel CNN-Transformer architecture that better attends to multi-scale features in images, making it more suitable for dense prediction tasks like image segmentation; (3) a sparse and dense feature fusion module to enhance collaboration between the two encoders. Experiments demonstrate that S2CCT outperforms previous state-of-the-art methods on two public medical image segmentation benchmarks, i.e., ACDC and KiTs19. The code and pretrained models will be released soon. Rongzhou Zhou, Ziqi Shu, Weixing Xie, Junfeng Yao, Qingqi Hong |
BIBM | 4 |
| 2024 | EmoTrans: Emotional Transition-based Model for Emotion Recognition in ConversationabstractIn an emotional conversation, emotions are causally transmitted among communication participants, constituting a fundamental conversational feature that can facilitate the comprehension of intricate changes in emotional states during the conversation and contribute to neutralizing emotional semantic bias in utterance caused by the absence of modality information. Therefore, emotional transition (ET) plays a crucial role in the task of Emotion Recognition in Conversation (ERC) that has not received sufficient attention in current research. In light of this, an Emotional Transition-based Emotion Recognizer (EmoTrans) is proposed in this paper. Specifically, we concatenate the most recent utterances with their corresponding speakers to construct the model input, known as samples, each with several placeholders to implicitly express the emotions of contextual utterances. Based on these placeholders, two components are developed to make the model sensitive to emotions and effectively capture the ET features in the sample. Furthermore, an ET-based Contrastive Learning (CL) is developed to compact the representation space, making the model achieve more robust sample representations. We conducted exhaustive experiments on four widely used datasets and obtained competitive experimental results, especially, new state-of-the-art results obtained on MELD and IEMOCAP, demonstrating the superiority of EmoTrans. Zhongquan Jian, Ante Wang, Jinsong Su, Junfeng Yao, Meihong Wang, Qingqiang Wu 0001 |
LREC/COLING | 4 |
| 2024 | GTPCR: Graph-Enhanced Transformer for Point Cloud RegistrationabstractAs Industry 4.0 continues to advance, point cloud registration technology is extensively employed in scenarios such as collaborative defect detection in products and digital twin-assisted assembly. In this paper, we present an end-to-end point cloud registration model, GTPCR, which is based on the utilization of spatial structural information within point cloud data. GTPCR employs a point-wise approach, directly solving rigid rotations based on estimated correspondences without reliance on RANSAC. It conceives of the point cloud as a graph structure in three-dimensional space, encoding nodes across various dimensions. The encoding of an individual node is dictated by its position and centrality. The geometric associations between nodes are contingent on factors such as relative position, Euclidean distance, azimuth, and elevation. Edge feature encoding is dynamically acquired through the learning process from node features. All encodings are incorporated as trainable biases input to the model. In comparison to methods transitioning from local to global, GTPCR demonstrates superior efficiency and heightened scalability. Moreover, due to its adept network architecture design, GTPCR effortlessly expands its applicability to non-rigid registration. Empirical evidence undeniably illustrates GTPCR’s competitive advantage across numerous datasets. In particular, GTPCR shows significant improvements in registration recall of 1.4% and 0.9% over baselines on the 3DMatch and 3DLoMatch datasets, respectively. Junfeng Yao, Yuanhang Li, Huabo Shen, Quan Qian, Xing Wu 0001 |
CSCWD | 2 |
| 2024 | A Collaborative Anomaly Localization Method Based on Multi-Modal ImagesabstractIn the context of industrial anomaly detection, anomaly point detection is a challenging task due to the rarity and unpredictable nature of anomalous samples. Existing 2D image-based defect detection methods have certain advantages in capturing features such as texture, color, and shape of parts. However, traditional single-modal defect detection methods (such as using only 2D images or only 3D point cloud data) may have limitations in accurately locating abnormal points when faced with complex surface defects on parts. Therefore, a collaborative abnormal localization method (CALM) based on multi-modal images is proposed to improve the accuracy of anomaly localization by fully utilizing information from multiple data sources. First, we propose a synchronized data augmentation method for 2D and 3D images to address the issue of scarce anomalous samples. Then, feature extraction is performed separately on RGB images and 3D point clouds, leveraging the features from both 2D and 3D images and performing multi-modal feature fusion while aligning the features. Finally, anomaly point localization and segmentation are achieved based on the abnormality scores output by the decoder. To validate the effectiveness of our method, experiments are conducted on the MVTec-3D AD dataset. The Pix-AUROC and Pix-AUPRO means of the CALM method reach 0.909 and 0.739, respectively. The experimental results demonstrate that our method achieves high detection accuracy at the pixel level, outperforming some traditional anomaly localization methods. Yuanhang Li, Junfeng Yao, Quan Qian, Xing Wu 0001 |
CSCWD | 2 |
| 2024 | Lightweight Transformer for sEMG Gesture Recognition with Feature Distilled Variational Information BottleneckabstractGesture recognition based on surface electromyography (sEMG) has seen considerable improvements in performance across various tasks and metrics with the rapid development of deep learning. However, challenges still exist in current deep neural networks for sEMG recognition. For instance, convolutional neural networks exhibit poor capturing of global features, recurrent neural networks have limited parallel processing capabilities, and their hybrids are usually more complex. Additionally, recent networks based on Transformers rarely consider the locality of attention and noise resistance. To fully explore the essence of sEMG sequences, and to make the model more lightweight and robust while ensuring feature learning performance, in this paper, we propose the feature distilled variational information bottleneck (FDVIB). Specifically, this method leverages knowledge distillation to learn from a high-precision teacher model at levels of feature and prediction, significantly reducing parameters and computations, and simplifying the structure. It also uses VIB to enhance the model’s robustness. We construct a Transformer model using the proposed method and conduct a series of evaluations. Experimental results show that our classification accuracy is competitive with state-of-the-art and also demonstrate the effectiveness of our method in enhancing model lightness and robustness. Junfeng Yao, Jinsong Su |
ECAI | 3 |
| 2024 | A Learning Rate Path Switching Training Paradigm for Version Updates of Large Language ModelsabstractDue to the continuous emergence of new data, version updates have become an indispensable requirement for Large Language Models (LLMs).The training paradigms for version updates of LLMs include pre-training from scratch (PTFS) and continual pre-training (CPT).Preliminary experiments demonstrate that PTFS achieves better pre-training performance, while CPT has lower training cost.Moreover, their performance and training cost gaps widen progressively with version updates.To investigate the underlying reasons for this phenomenon, we analyze the effect of learning rate adjustments during the two stages of CPT: preparing an initialization checkpoint and continual pre-training based on this checkpoint.We find that a large learning rate in the first stage and a complete learning rate decay process in the second stage are crucial for version updates of LLMs.Hence, we propose a learning rate path switching training paradigm.Our paradigm comprises one main path, where we pre-train a LLM with the maximal learning rate, and multiple branching paths, each of which corresponds to an update of the LLM with newly-added training data.Extensive experiments demonstrate the effectiveness and generalization of our paradigm.Particularly, when training four versions of LLMs, our paradigm reduces the total training cost to 58% compared to PTFS, while maintaining comparable pretraining performance. Jianheng Huang, Yixuan Liao, Xiaoxin Chen 0001, Junfeng Yao, Jinsong Su |
EMNLP | 7 |
| 2024 | SSFlowNet: Semi-supervised Scene Flow Estimation on Point Clouds With Pseudo Label
Jingze Chen, Simiao Zhuang, Qiqin Lin, Junfeng Yao |
ICANN (3) | 4 |
| 2024 | Conversation Clique-Based Model for Emotion Recognition In ConversationabstractEffective extraction and integration of valuable contextual information is the core of models for the Emotion Recognition in Conversation (ERC) task. However, a significant amount of irrelevant information is inevitably introduced when integrating long-range contextual information, perplexing the model greatly and resulting in incorrect emotion identification. To this end, we proposed a Conversation Clique-based Model (CCM), designed to extract the most efficacious contextual information to bolster the semantic quality of utterances. Specifically, we devise an utterance spatial relationship module (SpaRel) to explicitly model structural-level correlations among utterances by using GAT, and an emotion temporal relationship module (TemRel) to implicitly capture the emotion sequence constraints by employing HMM. We conduct extensive experiments on the publicly available MELD dataset, and the experimental results indicate the effectiveness of our proposed model, achieving new state-of-the-art results. Zhongquan Jian, Jiajian Li, Junfeng Yao, Meihong Wang, Qingqiang Wu 0001 |
ICASSP | 3 |
| 2024 | EDM: Synthetic Data from Exemplar Diffusion Model Improves Non-Communicable Diseases DetectionabstractThere have been researches revealing obvious associations between facial phenotypes and non-communicable diseases (NCDs), which enables effective health assessment with the integration of model-based learning methods. However, the paucity and poor quality of available datasets hinder the development of potent algorithms to detect NCDs. To meet this challenge, we propose a method called Exemplar Diffusion Model (EDM), the objective of proposed EDM is to generate facial images that illustrate simulated non-communicable diseases, utilizing a normal facial image as input. Extensive experimental results show that the proposed EDM method outperforms the state-of-the-art methods in terms of Frechet Inception Distance (FID) and Quality Score (QS), with improvements of 0.11 and 0.74, respectively. Furthermore, comprehensive ablation studies and comparative experiments prove the value of proposed EDM method in large-scale facial image dataset generation and non-communicable disease detection. Xing Wu 0001, Junfeng Yao, Quan Qian, Yike Guo |
ICASSP | 3 |
| 2024 | DRSM: Efficient Neural 4D Decomposition for Dynamic Reconstruction in Stationary Monocular CamerasabstractWith the popularity of monocular videos generated by video sharing and live broadcasting applications, reconstructing and editing dynamic scenes in stationary monocular cameras has become a special but anticipated technology. In contrast to scene reconstructions that exploit multi-view observations, the problem of modeling a dynamic scene from a single view is significantly more under-constrained and ill-posed. Inspired by recent progress in neural rendering, we present a novel framework to tackle 4D decomposition problem for dynamic scenes in monocular cameras. Our framework utilizes decomposed static and dynamic feature planes to represent 4D scenes and emphasizes the learning of dynamic regions through dense ray casting. Inadequate 3D clues from a single-view and occlusion are also particular challenges in scene reconstruction. To overcome these difficulties, we propose deep supervised optimization and ray casting strategies. With experiments on various videos, our method generates higher-fidelity results than existing methods for single-view dynamic scene representation. Weixing Xie, Qiqin Lin, Jingze Chen, Junfeng Yao, Xiaohu Guo |
ICASSP | 6 |
| 2024 | ICR-Net: Semi-Supervised Medical Image Segmentation Guided By Intra-Sample Cross ReconstructionabstractSemi-supervised learning is becoming increasingly popular in medical image segmentation because of its ability to exploit large amounts of unlabelled data to extract additional information. However, most existing semi-supervised segmentation methods focus only on extracting information from unlabelled data, ignoring the potential of labelled data to further improve model performance. In this paper, we propose a new framework for Intra-Sample Cross Reconstruction Networks (ICR-Net) that utilises labelled data to help the network extract information from unlabelled data, thereby guiding the network’s regularisation learning. Our method contains two modules: Intra-Sample Cross Reconstruction (ICR) module and Synergistic Consistency Constraints (SCC) module. The ICR module processes the labelled data features in a more fine-grained manner, thus enabling the network to learn and capture the key patterns and features in the inputs more efficiently, and the SCC guides the network’s regularised learning by formulating additional model regularisations. Experiments on the LA dataset and the pancreas dataset show that our proposed framework is more effective than current state-of-the-art methods in medical image segmentation tasks. Xianpeng Cao, Weixing Xie, Xianxing Cao, Qiqin Lin, Rongzhou Zhou, Junfeng Yao, Qingqi Hong |
ICME | 6 |
| 2024 | DPP-Net: Difficulty Perception-Processing Heterogeneous Network for Semi-supervised Medical Image SegmentationabstractIn semi-supervised medical image segmentation, the scarcity of labeled data makes models prone to learning bias, causing persistent errors in certain regions and eventual over-fitting, significantly impacting segmentation performance. These problematic regions, termed difficult areas, are inadequately addressed by existing methods. To address this, We propose the Difficulty Perception-Processing Heterogeneous Network (DPP-Net). It guides the model in accurately perceiving and rectifying difficult areas, overcoming learning bias. Specifically, we introduce the Global Mutual Perception (GMP) to establish a comprehensive information perception channel between sample data, enabling a more holistic and accurate perception of difficult areas. The Difficulty-Aware Rectification (DAR) structure ensures continuous monitoring of difficult areas during training, allowing for timely adjustments to errors. Additionally, the Adaptive Competitive Pseudo-Label (ACP) Augmentation strategy enhances pseudo-labels through adaptive confidence competition. Experimental results on two different medical image databases (CT and MRI) demonstrate that our approach outperforms several state-of-the-art methods. Qiqin Lin, Weixing Xie, Rongzhou Zhou, Xianpeng Cao, Jingze Chen, Junfeng Yao, Qingqi Hong |
ICME | 6 |
| 2024 | Audio-driven High-resolution Seamless Talking Head Video Editing via StyleGANabstractThe existing methods for audio-driven talking head video editing have the limitations of poor visual effects. This paper tries to tackle this problem through editing talking face images seamless with different emotions based on two modules: (1) an audio-to-landmark module, consisting of the CrossReconstructed Emotion Disentanglement and an alignment network module. It bridges the gap between speech and facial motions by predicting corresponding emotional landmarks from speech; (2) a landmark-based editing module edits face videos via StyleGAN. It aims to generate the seamless edited video consisting of the emotion and content components from the input audio. Extensive experiments confirm that compared with state-of-the-arts methods, our method provides high-resolution videos with high visual quality. Jiacheng Su, Kunhong Liu 0001, Junfeng Yao, Dongdong Lv |
ICME | 4 |
| 2024 | DMGCL: Denoising Multi-view Graph Contrastive Learning for Robust Recommendation
Xing Wu 0001, Mengkun Pi, Junfeng Yao, Quan Qian |
ICONIP (6) | 3 |
| 2024 | Hierarchical Adaptive Position Encoding-Based Transformer for Point Cloud Analysis
Jiawei Yao, Junfeng Yao, Chunyang Huang |
ICONIP (1) | 2 |
| 2024 | STMAE: Spatial Temporal Masked Auto-Encoder for Traffic Forecasting
Xing Wu 0001, Chengyou Cai, Jianjia Wang, Junfeng Yao, Quan Qian |
ICPR (5) | 5 |
| 2024 | Candidate Evaluation with Multimodal Data-Driven for Recruitment
Xing Wu 0001, Kehong Liu, Jianjia Wang, Junfeng Yao, Rongqi Lv |
ICPR (8) | 4 |
| 2024 | PaintingScapeVR: Exploring Gamified Interaction in Virtual Galleries
Dongning Cai, Le Ren, Yaqi Chang, Junfeng Yao |
ICXR | 6 |
| 2024 | SmartDeco: AR for Soft Furnishing Design
Yichun Zhang, Ruilong Liu, Shiyi Chen, Junfeng Yao |
ICXR | 6 |
| 2024 | SurgicalGaussian: Deformable 3D Gaussians for High-Fidelity Surgical Scene Reconstruction
Weixing Xie, Junfeng Yao, Xianpeng Cao, Qiqin Lin, Zerui Tang, Xiaohu Guo |
MICCAI (6) | 2 |
| 2024 | PGNET: A Real-Time Efficient Model for Underwater Object Detection
Hengsu Liu, Shibo Cong, Junfeng Yao |
PRCV (13) | 4 |
| 2024 | ICERS: Intelligent Collaborative Emergency Response SystemabstractAdvances in computer science and technology, coupled with the growing societal demand for enhanced safety, have paved the way for innovative smart home environments. While current smart home systems can send alarms and notify emergency contacts during a crisis, they do not provide timely and convenient on-site emergency assistance, potentially delaying critical aid and compromising user safety. This paper introduces the Intelligent Collaborative Emergency System (ICERS), which addresses this gap through three integrated subsystems: Fall Detection, Vital Signs Monitoring, and Central Dispatch. The Fall Detection module leverages camera video data to classify human positions and estimate potential falls. The Vital Signs Monitoring module employs millimeter-wave radar technology to provide non-contact, real-time health monitoring. The Central Dispatch system receives alert signals from on-edge devices and relays them to medical facilities and volunteer organizations for a coordinated emergency response. By utilizing multimodal data for collaborative efforts, the ICERS ensures that individuals receive timely notifications and coordinated emergency assistance during unforeseen situations. Junfeng Yao, Chengyou Cai, Yuelin Xu |
SoMeT | 2 |
| 2024 | A detail-preserving method for medial mesh computation in triangular meshesabstractThe medial axis transform (MAT) of an object is the set of all points inside the object that have more than one closest point on the object’s boundary. Representing sharp edges and corners of triangular meshes using MAT poses a complex challenge. While some researchers have proposed using zero-radius medial spheres to depict these features, they have not clearly articulated how to establish proper connections among them. In this paper, we propose a novel framework for computing MAT of a triangular mesh while preserving its features. The initial medial axis mesh obtained may contain erroneous edges, which are discussed and addressed in Section 3.3. Furthermore, during the simplification process, it is crucial to ensure that the medial spheres remain within the confines of the triangular mesh. Our algorithm excels in preserving critical features throughout the simplification procedure, consistently ensuring that the spheres remain enclosed within the triangular mesh. Experiments on various types of 3D models demonstrate the robustness, shape fidelity, and efficiency in representation achieved by our algorithm. Bingchuan Li, Yuping Ye, Junfeng Yao, Weixing Xie, Mengyuan Ge |
Graph. Model. | 3 |
| 2024 | FedEL: Federated ensemble learning for non-iid data
Xing Wu 0001, Jie Pei, Xianhua Han, Yen-Wei Chen 0001, Junfeng Yao, Yang Liu 0005, Quan Qian, Yike Guo |
Expert Syst. Appl. | 5 |
| 2024 | Retrieval Contrastive Learning for Aspect-Level Sentiment Classification
Zhongquan Jian, Jiajian Li, Qingqiang Wu 0001, Junfeng Yao |
Inf. Process. Manag. | 4 |
| 2024 | Transformer-based network with temporal depthwise convolutions for sEMG recognition
Junfeng Yao, Meiyan Xu, Min Jiang 0005, Jinsong Su |
Pattern Recognit. | 2 |
| 2024 | A Distance Transformation Deep Forest Framework With Hybrid-Feature Fusion for CXR Image ClassificationabstractDetecting pneumonia, especially coronavirus disease 2019 (COVID-19), from chest X-ray (CXR) images is one of the most effective ways for disease diagnosis and patient triage. The application of deep neural networks (DNNs) for CXR image classification is limited due to the small sample size of the well-curated data. To tackle this problem, this article proposes a distance transformation-based deep forest framework with hybrid-feature fusion (DTDF-HFF) for accurate CXR image classification. In our proposed method, hybrid features of CXR images are extracted in two ways: hand-crafted feature extraction and multigrained scanning. Different types of features are fed into different classifiers in the same layer of the deep forest (DF), and the prediction vector obtained at each layer is transformed to form distance vector based on a self-adaptive scheme. The distance vectors obtained by different classifiers are fused and concatenated with the original features, then input into the corresponding classifier at the next layer. The cascade grows until DTDF-HFF can no longer gain benefits from the new layer. We compare the proposed method with other methods on the public CXR datasets, and the experimental results show that the proposed method can achieve state-of-the art (SOTA) performance. The code will be made publicly available at https://github.com/hongqq/DTDF-HFF. Qingqi Hong, Lingli Lin, Qingde Li, Junfeng Yao, Qingqiang Wu 0001, Kunhong Liu 0001, Jie Tian 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Geometry-Based Molecular Generation With Deep Constrained Variational AutoencoderabstractFinding target molecules with specific chemical properties plays a decisive role in drug development. We proposed GEOM-CVAE, a constrained variational autoencoder based on geometric representation for molecular generation with specific properties, which is protein-context-dependent. In terms of machine learning, it includes continuous feature embedding encoder and molecular generation decoder. Our key contribution is to propose an efficient geometric embedding method, including the spatial structure representations of drug molecule (converting the 3-D coordinates into image) and the geometric graph representations of protein target (modeling the protein surface as a mesh). The 3-D geometric information is vital to successful molecular generation, which is different from previous molecular generative methods based on 1-D or 2-D. Our model framework generates specific molecules in two phases, by first generating special image with molecular 3-D information to learn latent representations and generating molecules with constrained condition based on geometric graph convolution for specific protein and then inputting the generated structural molecules into a parser network for obtaining Simplified Molecular Input Line Entry System (SMILES) strings. Our model achieves competitive performance that implies its potential effectiveness to enable the exploration of the vast chemical space for drug discovery. Chunyan Li 0002, Junfeng Yao, Wei Wei 0006, Zhangming Niu, Xiangxiang Zeng, Jin Li 0007, Jianmin Wang 0016 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Similar image matching via global topology consensusabstractAbstract Recovering three-dimensional structure from images is one of the important researches in computer vision. The quality of feature matching is one of the keys to obtaining more accurate results. However, as different objects or different surfaces of objects have similar images with the same elements and different typography, the camera pose estimation will be wrong and the task will fail. This paper proposes a new mismatch elimination algorithm based on global topology consistency. We first formulate the matching task as a mathematical model based on the global constraints, then convert the feature matching into grid matching, calculate the confidence of the grids according to the changes in the angle and displacement between correspondence grid vectors, and remove the mismatches with low confidence. The experiments have demonstrated that our proposed method performs better than the state-of-the-art feature matching methods to accomplish outlier match rejection in the task of similar image matching and could be helped to obtain the correct camera pose to reconstruct more complete and more accurate object models. Junfeng Yao, Junyi Long |
Vis. Comput. | 2 |
| 2024 | Automatic piano performance interaction system based on greedy algorithm for dexterous manipulatorabstractWith continuous advancements in artificial intelligence (AI), automatic piano-playing robots have become subjects of cross-disciplinary interest. However, in most studies, these robots served merely as objects of observation with limited user engagement or interaction. To address this issue, we propose a user-friendly and innovative interaction system based on the principles of greedy algorithms. This system features three modules: score management, performance control, and keyboard interactions. Upon importing a custom score or playing a note via an external device, the system performs on a virtual piano in line with user inputs. This system has been successfully integrated into our dexterous manipulator-based piano-playing device, which significantly enhances user interactions. Junfeng Yao, Yalan Zhou |
Virtual Real. Intell. Hardw. | 2 |
| 2023 | LagNet: Deep Lagrangian Mechanics for Plug-and-Play Molecular Representation LearningabstractMolecular representation learning is a fundamental problem in the field of drug discovery and molecular science. Whereas incorporating molecular 3D information in the representations of molecule seems beneficial, which is related to computational chemistry with the basic task of predicting stable 3D structures (conformations) of molecules. Existing machine learning methods either rely on 1D and 2D molecular properties or simulate molecular force field to use additional 3D structure information via Hamiltonian network. The former has the disadvantage of ignoring important 3D structure features, while the latter has the disadvantage that existing Hamiltonian neural network must satisfy the “canonial” constraint, which is difficult to be obeyed in many cases. In this paper, we propose a novel plug-and-play architecture LagNet by simulating molecular force field only with parameterized position coordinates, which implements Lagrangian mechanics to learn molecular representation by preserving 3D conformation without obeying any additional restrictions. LagNet is designed to generate known conformations and generalize for unknown ones from molecular SMILES. Implicit positions in LagNet are learned iteratively using discrete-time Lagrangian equations. Experimental results show that LagNet can well learn 3D molecular structure features, and outperforms previous state-of-the-art baselines related molecular representation by a significant margin. Chunyan Li 0002, Junfeng Yao, Jinsong Su, Zhaoyang Liu 0002, Xiangxiang Zeng |
AAAI | 2 |
| 2023 | Knowledge Graph Embedding with Relation Rotation and Entity Adjustment by Quaternions
Wen Sun 0007, Qingqiang Wu 0001, Xiaoli Wang 0002, Junfeng Yao, Zhifeng Bao |
ADMA (4) | 4 |
| 2023 | MEA-TransUNet: A Multiple External Attention Network for Multi-Organ Segmentation
Xianpeng Cao, Junfeng Yao, Qingqi Hong, Rongzhou Zhou |
ICANN (9) | 2 |
| 2023 | Micro-expression Recognition Based on PCB-PCANet+
Shiqi Wang 0020, Junfeng Yao |
ICONIP (6) | 3 |
| 2023 | End-to-End Optical Music Recognition with Attention Mechanism and Memory Units Optimization
Ruichen He, Junfeng Yao |
PRCV (2) | 2 |
| 2023 | Cross Attention Multi Scale CNN-Transformer Hybrid Encoder Is General Medical Image Learner
Rongzhou Zhou, Junfeng Yao, Qingqi Hong, Xingxin Li, Xianpeng Cao |
PRCV (13) | 2 |
| 2023 | LATrans-Unet: Improving CNN-Transformer with Location Adaptive for Medical Image Segmentation
Qiqin Lin, Junfeng Yao, Qingqi Hong, Xianpeng Cao, Rongzhou Zhou, Weixing Xie |
PRCV (13) | 2 |
| 2023 | Combining Structure Embedding and Text Semantics for Efficient Knowledge Graph CompletionabstractKnowledge graph completion plays a crucial role in downstream applications.However, existing methods tend to only rely on the structure or textual information, resulting in suboptimal model performance.Moreover, recent attempts to leverage pre-trained language models to complete knowledge graphs have proved unsatisfactory.To overcome these limitations, we propose a novel model that combines structural embedding and semantic information of the knowledge graph.Compared with previous works based on pre-trained language models, our model can better use the implicit knowledge of pre-trained language models by using relation templates, entity definitions, and learnable tokens.Furthermore, our model employs a multihead attention mechanism to transform the embedding semantic space of entities and relations obtained from the knowledge graph embedding model, thereby enhancing their expressiveness and unifying the semantic space of both types of information.Finally, we utilize convolutional neural networks to extract features from the matrices created by combining these two types of information for link prediction and triplet classification tasks.Empirical evaluations on two knowledge graph completion datasets demonstrate that our model is effective for both tasks. Wen Sun 0007, Junfeng Yao, Qingqiang Wu 0001, Kunhong Liu 0001 |
SEKE | 3 |
| 2023 | Multi-Layer Transformer for Video ClassificationabstractVideo classification is a challenging task because of the intricate spatiotemporal information present within videos. Current models often rely on 2D or 3D convolutional neural networks. However, convolutional neural networks are difficult to solve the long-range dependency problem. In addition, they are computationally expensive and memory-intensive. To address the challenges, a Multi-layer Transformer is proposed for video classification. The proposed method takes advantage of the high correlation between adjacent frames by grouping them and learning local and global information with a multi-layer structure based on Transformer. First, different frame sampling rates and grouping strategies are tested in the experiments, then comparing the method with state-of-the-art models. The results demonstrate that the proposed method has advanced performance with TOP1 accuracy of 77.8% on the Kinetics-400 dataset and 64.9% on the Something-Something v2 dataset. Xing Wu 0001, Chenjie Tao, Junfeng Yao, Quan Qian |
SoMeT | 3 |
| 2023 | Point2MM: Learning medial mesh from point clouds
Mengyuan Ge, Junfeng Yao, Baorong Yang, Ningna Wang, Zhonggui Chen, Xiaohu Guo |
Comput. Graph. | 2 |
| 2023 | Efficient collision detection using hybrid medial axis transform and BVH for rigid body simulationabstractMedial Axis Transform (MAT) has been recently adopted as the acceleration structure of broad-phase collision detection. Compared to traditional BVH-based methods, MAT can provide a high-fidelity volumetric approximation of 3D complex objects, resulting in higher collision culling efficiency. However, due to MAT’s non-hierarchical structure, it may be outperformed in collision-light scenarios because several cullings at the top level of a BVH may take a large number of cullings with MAT. We propose a collision detection method that combines MAT and BVH to address the above problem. Our technique efficiently culls collisions between dynamic and static objects. Experimental results show that our method has higher culling efficiency than pure BVH or MAT methods. Xingxin Li, Shibo Song, Junfeng Yao, Hanyin Zhang, Rongzhou Zhou, Qingqi Hong |
Graph. Model. | 3 |
| 2023 | A self-adaptive soft-recoding strategy for performance improvement of error-correcting output codes
Guangyi Lin, Nan Zeng, Yong Xu 0009, Kunhong Liu 0001, Beizhan Wang, Junfeng Yao, Qingqiang Wu 0001 |
Pattern Recognit. | 7 |
| 2023 | D$^{2}$PSG: Multi-Party Dialogue Discourse Parsing as Sequence GenerationabstractConversational discourse analysis aims to extract the interactions between dialogue turns, which is crucial for modeling complex multi-party dialogues. As the benchmarks are still limited in size and human annotations are costly, the current standard approaches apply pretrained language models, but they still require randomly initialized classifiers to make predictions. These classifiers usually require massive data to work smoothly with the pretrained encoder, causing severe data hunger issue. We propose two convenient strategies to formulate this task as a sequence generation problem, where classifier decisions are carefully converted into sequence of tokens. We then adopt a pretrained T5 1 model to solve this task so that no parameters are randomly initialized. We also leverage the descriptions of the discourse relations to help model understand their meanings. Experiments on two popular benchmarks show that our approach outperforms previous state-of-the-art models by a large margin, and it is also more robust in zero-shot and few-shot settings. Ante Wang, Linfeng Song, Lifeng Jin, Junfeng Yao, Haitao Mi, Chen Lin 0001, Jinsong Su, Dong Yu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Block Division Convolutional Network With Implicit Deep Features Augmentation for Micro-Expression RecognitionabstractDespite the development of computer vision techniques, the micro-expression (ME) recognition task still remains a great challenge because MEs have very low intensity and short duration. However, the ME recognition is of great significance since it provides important clues for real affective states detection. This paper proposes a novel Block Division Convolutional Network (BDCNN) with the implicit deep features augmentation. In detail, BDCNN learns from four optical flow features computed by the onset and apex frames of each video. It innovatively divides each image into a set of small blocks in the deep learning model, then the convolution and pooling operations are performed on these small blocks in sequence. To handle the small sample size problem in the micro-expression data, this study uses the improved implicit semantic data augmentation algorithm in the deep features space. Experiments are conducted on three publicly available databases, viz, CASME II, SMIC, and SAMM. Experimental results show that our model outperforms the state-of-the-art methods by attaining the accuracy of 84.32% and F1-score of 82.13% on the 3-class datasets, and the accuracy of 81.82% and F1-score of 75.46% on the 5-class datasets, respectively. Our source code is publicly available for non-commercial or research use athttps://github.com/MLDMXM2017/BDCNN. Bin Chen 0024, Kunhong Liu 0001, Yong Xu 0009, Qingqiang Wu 0001, Junfeng Yao |
IEEE Trans. Multim. | 5 |
| 2023 | Inverse transformation sampling-based attentive cutout for fine-grained visual recognition
Yaojin Lin, Meiyan Xu, Ming-Wen Shao, Junfeng Yao |
Vis. Comput. | 5 |
| 2023 | Outliers rejection in similar image matchingabstractImage matching is crucial in numerous computer vision tasks such as 3D reconstruction and simultaneous visual localization and mapping. The accuracy of the matching significantly impacted subsequent studies. Because of their local similarity, when image pairs contain comparable patterns but feature pairs are positioned differently, incorrect recognition can occur as global motion consistency is disregarded. This study proposes an image-matching filtering algorithm based on global motion consistency. It can be used as a subsequent matching filter for the initial matching results generated by other matching algorithms based on the principle of motion smoothness. A particular matching algorithm can first be used to perform the initial matching; then, the rotation and movement information of the global feature vectors are combined to effectively identify outlier matches. The principle is that if the matching result is accurate, the feature vectors formed by any matched point should have similar rotation angles and moving distances. Thus, global motion direction and global motion distance consistencies were used to reject outliers caused by similar patterns in different locations. Four datasets were used to test the effectiveness of the proposed method. Three datasets with similar patterns in different locations were used to test the results for similar images that could easily be incorrectly matched by other algorithms, and one commonly used dataset was used to test the results for the general image-matching problem. The experimental results suggest that the proposed method is more accurate than other state-of-the-art algorithms in identifying mismatches in the initial matching set. The proposed outlier rejection matching method can significantly improve the matching accuracy for similar images with locally similar feature pairs in different locations and can provide more accurate matching results for subsequent computer vision tasks. Junfeng Yao |
Virtual Real. Intell. Hardw. | 2 |
| 2022 | The design of error-correcting output codes algorithm for the open-set recognition
Kunhong Liu 0001, Zhan WangPing, Yi-Fan Liang, Hong-Zhou Guo, Junfeng Yao, Qingqiang Wu 0001, Qingqi Hong |
Appl. Intell. | 6 |
| 2022 | Effective knowledge graph embeddings based on multidirectional semantics relations for polypharmacy side effects predictionabstractMOTIVATION: Polypharmacy is the combined use of drugs for the treatment of diseases. However, it often shows a high risk of side effects. Due to unnecessary interactions of combined drugs, the side effects of polypharmacy increase the risk of disease and even lead to death. Thus, obtaining abundant and comprehensive information on the side effects of polypharmacy is a vital task in the healthcare industry. Early traditional methods used machine learning techniques to predict side effects. However, they often make costly efforts to extract features of drugs for prediction. Later, several methods based on knowledge graphs are proposed. They are reported to outperform traditional methods. However, they still show limited performance by failing to model complex relations of side effects among drugs. RESULTS: To resolve the above problems, we propose a novel model by further incorporating complex relations of side effects into knowledge graph embeddings. Our model can translate and transmit multidirectional semantics with fewer parameters, leading to better scalability in large-scale knowledge graphs. Experimental evaluation shows that our model outperforms state-of-the-art models in terms of the average area under the ROC and precision-recall curves. AVAILABILITY AND IMPLEMENTATION: Code and data are available at: https://github.com/galaxysunwen/MSTE-master. Junfeng Yao, Wen Sun 0007, Zhongquan Jian, Qingqiang Wu 0001, Xiaoli Wang 0002 |
Bioinform. | 1 |
| 2022 | GPU-based supervoxel segmentation for 3D point clouds
Yanyang Xiao, Zhonggui Chen, Junfeng Yao, Xiaohu Guo |
Comput. Aided Geom. Des. | 4 |
| 2022 | FTAP: Feature transferring autonomous machine learning pipeline
Xing Wu 0001, Cheng Chen 0075, Mingyu Zhong, Jianjia Wang, Quan Qian, Junfeng Yao, Yike Guo |
Inf. Sci. | 8 |
| 2022 | AAN+: Generalized Average Attention Network for Accelerating Neural TransformerabstractTransformer benefits from the high parallelization of attention networks in fast training, but it still suffers from slow decoding partially due to the linear dependency O(m) of the decoder self-attention on previous target words at inference. In this paper, we propose a generalized average attention network (AAN+) aiming at speeding up decoding by reducing the dependency from O(m) to O(1). We find that the learned self-attention weights in the decoder follow some patterns which can be approximated via a dynamic structure. Based on this insight, we develop AAN+, extending our previously proposed average attention (Zhang et al., 2018a, AAN) to support more general position- and content-based attention patterns. AAN+ only requires to maintain a small constant number of hidden states during decoding, ensuring its O(1) dependency. We apply AAN+ as a drop-in replacement of the decoder selfattention and conduct experiments on machine translation (with diverse language pairs), table-to-text generation and document summarization. With masking tricks and dynamic programming, AAN+ enables Transformer to decode sentences around 20% faster without largely compromising in the training speed and the generation performance. Our results further reveal the importance of the localness (neighboring words) in AAN+ and its capability in modeling long-range dependency. Biao Zhang 0002, Deyi Xiong, Yubin Ge, Junfeng Yao, Jinsong Su |
J. Artif. Intell. Res. | 4 |
| 2022 | An AST Structure Enhanced Decoder for Code GenerationabstractCurrently, the most dominant neural code generation modelsare often equipped with a tree-structured LSTM decoder, which outputs a sequence of actions to construct an Abstract Syntax Tree (AST) via pre-order traversal. However, such a decoder has two obvious drawbacks. First, except for the parent action, other faraway and important history actions rarely contribute to the current decision. Second, it also neglects future actions, which may be crucial for the prediction of the current action. To deal with these issues, in this paper, we propose a novel AST structure enhanced decoder for code generation, which significantly extends the decoder with respect to the above two aspects. First, we introduce an AST information enhanced attention mechanism to fully exploit history actions, of which impacts are further distinguished according to their syntactic distances, action types and relative positions; Second, we jointly model the predictions of current action and its important future action via multi-task learning, where the learned hidden state of the latter can be further leveraged to improve the former. Experimental results on commonly-used datasets demonstrate the effectiveness of our proposed decoder.1 Linfeng Song, Yubin Ge, Fandong Meng, Junfeng Yao, Jinsong Su |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | 3DMol-Net: Learn 3D Molecular Representation Using Adaptive Graph Convolutional Network Based on Rotation InvarianceabstractStudying the deep learning-based molecular representation has great significance on predicting molecular property, promoted the development of drug screening and new drug discovery, and improving human well-being for avoiding illnesses. It is essential to learn the characterization of drug for various downstream tasks, such as molecular property prediction. In particular, the 3D structure features of molecules play an important role in biochemical function and activity prediction. The 3D characteristics of molecules largely determine the properties of the drug and the binding characteristics of the target. However, most current methods merely rely on 1D or 2D properties while ignoring the 3D topological structure, thereby degrading the performance of molecular inferring. In this paper, we propose 3DMol-Net to enhance the molecular representation, considering both the topology and rotation invariance (RI) of the 3D molecular structure. Specifically, we construct a molecular graph with soft relations related to the spatial arrangement of the 3D coordinates to learn 3D topology of arbitrary graph structure and employ an adaptive graph convolutional network to predict molecular properties and biochemical activities. Comparing with current graph-based methods, 3DMol-Net demonstrates superior performance in terms of both regression and classification tasks. Further verification of RI and visualization also show better robustness and representation capacity of our model. Chunyan Li 0002, Wei Wei 0006, Jin Li 0007, Junfeng Yao, Xiangxiang Zeng, Zhihan Lyu |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | Improving Tree-Structured Decoder Training for Code Generation via Mutual LearningabstractCode generation aims to automatically generate a piece of code given an input natural language utterance. Currently, among dominant models, it is treated as a sequence-to-tree task, where a decoder outputs a sequence of actions corresponding to the pre-order traversal of an Abstract Syntax Tree. However, such a decoder only exploits the pre-order traversal based preceding actions, which are insufficient to ensure correct action predictions. In this paper, we first throughly analyze the context modeling difference between neural code generation models with different traversals based decodings (preorder traversal vs breadth-first traversal), and then propose to introduce a mutual learning framework to jointly train these models. Under this framework, we continuously enhance both two models via mutual distillation, which involves synchronous executions of two one-to-one knowledge transfers at each training step. More specifically, we alternately choose one model as the student and the other as its teacher, and require the student to fit the training data and the action prediction distributions of its teacher. By doing so, both models can fully absorb the knowledge from each other and thus could be improved simultaneously. Experimental results and in-depth analysis on several benchmark datasets demonstrate the effectiveness of our approach. We release our code at https://github.com/DeepLearnXMU/CGML. Binbin Xie, Jinsong Su, Yubin Ge, Xiang Li 0104, Jianwei Cui 0002, Junfeng Yao, Bin Wang 0004 |
AAAI | 6 |
| 2021 | Improving Graph-based Sentence Ordering with Iteratively Predicted Pairwise OrderingsabstractShaopeng Lai, Ante Wang, Fandong Meng, Jie Zhou, Yubin Ge, Jiali Zeng, Junfeng Yao, Degen Huang, Jinsong Su. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Shaopeng Lai, Ante Wang, Fandong Meng, Jie Zhou 0016, Yubin Ge, Jiali Zeng, Junfeng Yao, Degen Huang, Jinsong Su |
EMNLP (1) | 7 |
| 2021 | A Structure Self-Aware Model for Discourse Parsing on Multi-Party DialoguesabstractConversational discourse structures aim to describe how a dialogue is organized, thus they are helpful for dialogue understanding and response generation. This paper focuses on predicting discourse dependency structures for multi-party dialogues. Previous work adopts incremental methods that take the features from the already predicted discourse relations to help generate the next one. Although the inter-correlations among predictions considered, we find that the error propagation is also very serious and hurts the overall performance. To alleviate error propagation, we propose a Structure Self-Aware (SSA) model, which adopts a novel edge-centric Graph Neural Network (GNN) to update the information between each Elementary Discourse Unit (EDU) pair layer by layer, so that expressive representations can be learned without historical predictions. In addition, we take auxiliary training signals (e.g. structure distillation) for better representation learning. Our model achieves the new state-of-the-art performances on two conversational discourse parsing benchmarks, largely outperforming the previous methods. Ante Wang, Linfeng Song, Shaopeng Lai, Junfeng Yao, Jinsong Su |
IJCAI | 5 |
| 2021 | A spatial-temporal gated attention module for molecular property prediction based on molecular geometryabstractMOTIVATION: Geometry-based properties and characteristics of drug molecules play an important role in drug development for virtual screening in computational chemistry. The 3D characteristics of molecules largely determine the properties of the drug and the binding characteristics of the target. However, most of the previous studies focused on 1D or 2D molecular descriptors while ignoring the 3D topological structure, thereby degrading the performance of molecule-related prediction. Because it is very time-consuming to use dynamics to simulate molecular 3D conformer, we aim to use machine learning to represent 3D molecules by using the generated 3D molecular coordinates from the 2D structure. RESULTS: We proposed Drug3D-Net, a novel deep neural network architecture based on the spatial geometric structure of molecules for predicting molecular properties. It is grid-based 3D convolutional neural network with spatial-temporal gated attention module, which can extract the geometric features for molecular prediction tasks in the process of convolution. The effectiveness of Drug3D-Net is verified on the public molecular datasets. Compared with other deep learning methods, Drug3D-Net shows superior performance in predicting molecular properties and biochemical activities. AVAILABILITY AND IMPLEMENTATION: https://github.com/anny0316/Drug3D-Net. SUPPLEMENTARY DATA: Supplementary data are available online at https://academic.oup.com/bib. Chunyan Li 0002, Jianmin Wang 0016, Zhangming Niu, Junfeng Yao, Xiangxiang Zeng |
Briefings Bioinform. | 4 |
| 2021 | An Efficient Tongue Segmentation Model Based on U-Net FrameworkabstractIn the intelligently processing of the tongue image, one of the most important tasks is to accurately segment the tongue body from a whole tongue image, and the good quality of tongue body edge processing is of great significance for the relevant tongue feature extraction. To improve the performance of the segmentation model for tongue images, we propose an efficient tongue segmentation model based on U-Net. Three important studies are launched, including optimizing the model’s main network, innovating a new network to specially handle tongue edge cutting and proposing a weighted binary cross-entropy loss function. The purpose of optimizing the tongue image main segmentation network is to make the model recognize the foreground and background features for the tongue image as well as possible. A novel tongue edge segmentation network is used to focus on handling the tongue edge because the edge of the tongue contains a number of important information. Furthermore, the advantageous loss function proposed is to be adopted to enhance the pixel supervision corresponding to tongue images. Moreover, thanks to a lack of tongue image resources on Traditional Chinese Medicine (TCM), some special measures are adopted to augment training samples. Various comparing experiments on two datasets were conducted to verify the performance of the segmentation model. The experimental results indicate that the loss rate of our model converges faster than the others. It is proved that our model has better stability and robustness of segmentation for tongue image from poor environment. The experimental results also indicate that our model outperforms the state-of-the-art ones in aspects of the two most important tongue image segmentation indexes: IoU and Dice. Moreover, experimental results on augmentation samples demonstrate our model have better performances. Qunsheng Ruan, Junfeng Yao, Yingdong Wang, Hsien-Wei Tseng, Zhiling Zhang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2021 | An External Knowledge Enhanced Graph-based Neural Network for Sentence OrderingabstractAs an important text coherence modeling task, sentence ordering aims to coherently organize a given set of unordered sentences. To achieve this goal, the most important step is to effectively capture and exploit global dependencies among these sentences. In this paper, we propose a novel and flexible external knowledge enhanced graph-based neural network for sentence ordering. Specifically, we first represent the input sentences as a graph, where various kinds of relations (i.e., entity-entity, sentence-sentence and entity-sentence) are exploited to make the graph representation more expressive and less noisy. Then, we introduce graph recurrent network to learn semantic representations of the sentences. To demonstrate the effectiveness of our model, we conduct experiments on several benchmark datasets. The experimental results and in-depth analysis show our model significantly outperforms the existing state-of-the-art models. Yongjing Yin, Shaopeng Lai, Linfeng Song, Chulun Zhou, Xianpei Han, Junfeng Yao, Jinsong Su |
J. Artif. Intell. Res. | 6 |
| 2021 | Multi-Rhythm Capsule Network Recognition Structure for Motor Imagery ClassificationabstractExisting machine learning methods for classification and recognition of EEG motor imagery usually suffer from reduced accuracy for limited training data. To address this problem, this paper proposes a multi-rhythm capsule network (FBCapsNet) that uses as little EEG information as possible with key features to classify motor imagery and further improves the classification efficiency. The network conforms to a small recognition model with only 3 acquisition channels but it can effectively use the limited data for feature learning. Based on the BCI Competition IV 2b data set, experimental results show that the proposed network can achieve 2.41% better performance than existing cutting-edge methods. Meiyan Xu, Junfeng Yao, Yifeng Zheng 0004, Yaojin Lin |
J. Web Eng. | 2 |
| 2021 | GPU-Based Supervoxel Generation With a Novel Anisotropic MetricabstractVideo over-segmentation into supervoxels is an important pre-processing technique for many computer vision tasks. Videos are an order of magnitude larger than images. Most existing methods for generating supervovels are either memory- or time-inefficient, which limits their application in subsequent video processing tasks. In this paper, we present an anisotropic supervoxel method, which is memory-efficient and can be executed on the graphics processing unit (GPU). Therefore, our algorithm achieves good balance among segmentation quality, memory usage and processing time. In order to provide accurate segmentation for moving objects in video, we use the optical flow information to design a brand new non-Euclidean metric to calculate the anisotropic distances between seeds and voxels. To efficiently compute the anisotropic metric, we adjust the classic jump flooding algorithm (which is designed for parallel execution on the GPU) to generate anisotropic Voronoi tessellation in the combined color and spatio-temporal space. We evaluate our method and the representative supervoxel algorithms for their capability on segmentation performance, computation speed and memory efficiency. We also apply supervoxel results to the application of foreground propagation in videos to test the performance on solving practical problems. Experiments show that our algorithm is much faster than the existing methods, and achieves good balance on segmentation quality and efficiency. Zhonggui Chen, Yong-Jin Liu 0001, Junfeng Yao, Xiaohu Guo |
IEEE Trans. Image Process. | 4 |
| 2021 | Medial IPC: accelerated incremental potential contact with medial elasticsabstractWe propose a framework of efficient nonlinear deformable simulation with both fast continuous collision detection and robust collision resolution. We name this new framework Medial IPC as it integrates the merits from medial elastics, for an efficient and versatile reduced simulation, as well as incremental potential contact, for a robust collision and contact resolution. We leverage medial axis transform to construct a kinematic subspace. Instead of resorting to projective dynamics, we use classic hyperelastics to embrace real-world nonlinear materials. A novel reduced continuous collision detection algorithm is presented based on the medial mesh. Thanks to unique geometric properties of medial axis and medial primitives, we derive closed-form formulations for identifying between-primitive collision within the reduced medial space. In the meantime, the implicit barrier energy that generates necessary repulsion forces for collision resolution is also formulated with the medial coordinate. In other words, Medial IPC exploits a universal reduced coordinate for simulation, continuous self-/collision detection, and IPC-based collision resolution. Continuous collision detection also allows more aggressive time stepping. In addition, we carefully implement our system with a heterogeneous CPU-GPU deployment such that massively parallelizable computations are carried out on the GPU while few sequential computations are on the CPU. Such implementation also frees us from generating training poses for selecting Cubature points and pre-computing their weights. We have tested our method on complicated deformable models and collision-rich simulation scenarios. Due to the reduced nature of our system, the computation is faster than fullspace IPC or other fullspace methods using continuous collision detection by at least one order. The simulation remains high-quality as the medial subspace captures intriguing and local deformations with sufficient realism. Lei Lan, Yin Yang 0002, Danny M. Kaufman, Junfeng Yao, Minchen Li, Chenfanfu Jiang |
ACM Trans. Graph. | 4 |
| 2020 | P2MAT-NET: Learning medial axis transform from sparse point clouds
Baorong Yang, Junfeng Yao, Bin Wang 0021, Jianwei Hu 0003, Yiling Pan, Tianxiang Pan, Wenping Wang 0001, Xiaohu Guo |
Comput. Aided Geom. Des. | 2 |
| 2020 | Learning EEG topographical representation for classification via convolutional neural network
Meiyan Xu, Junfeng Yao, Zhihong Zhang 0001, Baorong Yang, Chunyan Li 0002, Junsong Zhang |
Pattern Recognit. | 2 |
| 2020 | Medial Elastics: Efficient and Collision-Ready Deformation via Medial Axis TransformabstractWe propose a framework for the interactive simulation of nonlinear deformable objects. The primary feature of our system is the seamless integration of deformable simulation and collision culling, which are often independently handled in existing animation systems. The bridge connecting them is the medial axis transform (MAT), a high-fidelity volumetric approximation of complex 3D shapes. From the physics simulation perspective, MAT leads to an expressive and compact reduced nonlinear model. We employ a semireduced projective dynamics formulation, which well captures high-frequency local deformations of high-resolution models while retaining a low computation cost. Our key observation is that the most compelling (nonlinear) deformable effects are enabled by the local constraints projection, which should not be aggressively reduced, and only apply model reduction at the global stage. From the collision detection (CD)/collision culling (CC) perspective, MAT is geometrically versatile using linear-interpolated spheres (i.e., the so-called medial primitives (MPs)) to approximate the boundary of the input model. The intersection test between two MPs is formulated as a quadratically constrained quadratic program problem. We give an algorithm to solve this problem exactly, which returns the deepest penetration between a pair of intersecting MPs. When coupled with spatial hashing, collision (including self-collision) can be efficiently identified on the GPU within a few milliseconds even for massive simulations. We have tested our system on a variety of geometrically complex and high-resolution deformable objects, and our system produces convincing animations with all of the collisions/self-collisions well handled at an interactive rate. Lei Lan, Ran Luo 0001, Marco Fratarcangeli, Weiwei Xu 0003, Huamin Wang 0001, Xiaohu Guo, Junfeng Yao, Yin Yang 0002 |
ACM Trans. Graph. | 7 |
| 2019 | Application and Research of Image-Based Modeling and 3D Printing Technology in Intangible Cultural Heritage Quanzhou Marionette ProtectionabstractWe present a practical solution to the modeling improvements for Quanzhou Marionette in 3D printing. We use 3D printing technology to improve the marionette production process. We have completed rapid image-based modeling to aid in design. According to the actual situation, the special parts such as marionette heads and joints are modeled and improved. The experiment proves the advantages of 3D printing in marionette production. 3D printing can be used as a production tool method for the protection and inheritance of intangible cultural heritage. Junfeng Yao, Kaini Huang |
COMPSAC (1) | 2 |
| 2019 | Exploiting reverse target-side contexts for neural machine translation via asynchronous bidirectional decoding
Jinsong Su, Xiangwen Zhang, Junfeng Yao, Yang Liu 0005 |
Artif. Intell. | 5 |
| 2019 | Superpixel Generation by Agglomerative Clustering With Quadratic Error MinimizationabstractAbstract Superpixel segmentation is a popular image pre‐processing technique in many computer vision applications. In this paper, we present a novel superpixel generation algorithm by agglomerative clustering with quadratic error minimization. We use a quadratic error metric (QEM) to measure the difference of spatial compactness and colour homogeneity between superpixels. Based on the quadratic function, we propose a bottom‐up greedy clustering algorithm to obtain higher quality superpixel segmentation. There are two steps in our algorithm: merging and swapping. First, we calculate the merging cost of two superpixels and iteratively merge the pair with the minimum cost until the termination condition is satisfied. Then, we optimize the boundary of superpixels by swapping pixels according to their swapping cost to improve the compactness. Due to the quadratic nature of the energy function, each of these atomic operations has only O(1) time complexity. We compare the new method with other state‐of‐the‐art superpixel generation algorithms on two datasets, and our algorithm demonstrates superior performance. Zhonggui Chen, Junfeng Yao, Xiaohu Guo |
Comput. Graph. Forum | 3 |
| 2018 | Real-Time Eye-Gaze Based Interaction for Human Intention Prediction and Emotion AnalysisabstractThe human eye's state of motion and content of interest can express people's cognitive status and emotional status based on their situation. When observing the surrounding things, the human eyes make different eye movements according to the observed objects which reflects human's attention and interest. In this paper, we capture and analyze patterns of human eye-gaze behavior and head motion and classify them into different categories. Besides, we compute and train the eye-object movement attention model and eye-object feature preference model based on different peoples' eye-gaze behaviors by using machine learning algorithms. These models are used to predict humans' object of interest and the interaction intention according to people's real-time situation. Furthermore, the eye-gaze behavior and head motion patterns can be used as a modality of non-verbal information in the computing of human emotional states based on the PAD affective computing model. Our methodology analyzes human emotion and cognition status from the aspect of eye-gaze behavior and head motion, understands the cognitive information that human eyes can express, and effectively improves the efficiency of human-computer interaction in different circumstances. Yingying She, Jianbing Xiahou, Junfeng Yao, Jun Li 0043, Qingqi Hong, Yingxuan Ji |
CGI | 4 |
| 2018 | DMAT: Deformable Medial Axis Transform for Animated Mesh ApproximationabstractAbstract Extracting a faithful and compact representation of an animated surface mesh is an important problem for computer graphics. However, the surface‐based methods have limited approximation power for volume preservation when the animated sequences are extremely simplified. In this paper, we introduce Deformable Medial Axis Transform (DMAT), which is deformable medial mesh composed of a set of animated spheres. Starting from extracting an accurate and compact representation of a static MAT as the template and partitioning the vertices on the input surface as the correspondences for each medial primitive, we present a correspondence‐based approximation method equipped with an As‐Rigid‐As‐Possible (ARAP) deformation energy defined on medial primitives. As a result, our algorithm produces DMAT with consistent connectivity across the whole sequence, accurately approximating the input animated surfaces. Baorong Yang, Junfeng Yao, Xiaohu Guo |
Comput. Graph. Forum | 2 |
| 2017 | Learning Spatiotemporal and Geometric Features with ISA for Video-Based Facial Expression Recognition
ChenHan Lin, Junfeng Yao, Ming-Ting Sun, Jinsong Su |
ICONIP (3) | 3 |
| 2017 | A Feature Learning Approach for Image Retrieval
Junfeng Yao, Yao Yu 0003, Yukai Deng, Changyin Sun 0001 |
ICONIP (2) | 1 |
| 2017 | Texturing of augmented reality character based on colored drawingabstractColoring book can inspire imaginary and creativity of children. However, with the rapid development of digital devices and internet, traditional coloring book tends to be not attractive for children any more. Thus, we propose an idea of applying augmented reality technology to traditional coloring book. After children finish coloring characters in the printed coloring book, they can inspect their work using a mobile device. The drawing is detected and tracked so that the video stream is augmented with a 3D character textured according to their coloring. This is possible thanks to several novel technical contributions. We present a texture process that generates texture map for 3D augmented reality character from 2D colored drawing using a lookup map. Considering the movement of the mobile device and drawing, we give an efficient method to track the drawing surface. Hengheng Zhao, Junfeng Yao |
VR | 3 |
| 2017 | Customization and fabrication of the appearance for humanoid robot
Shihui Guo, Hanxiang Xu, Nadia Magnenat-Thalmann, Junfeng Yao |
Vis. Comput. | 4 |
| 2017 | Medial-axis-driven shape deformation with volume preservation
Lei Lan, Junfeng Yao, Xiaohu Guo |
Vis. Comput. | 2 |
| 2016 | An implicit skeleton-based method for the geometry reconstruction of vasculatures
Qingqi Hong, Qingde Li, Beizhan Wang, Junfeng Yao, Qingqiang Wu 0001, Yingying She |
Vis. Comput. | 5 |
| 2015 | A Context-Aware Topic Model for Statistical Machine TranslationabstractJinsong Su, Deyi Xiong, Yang Liu, Xianpei Han, Hongyu Lin, Junfeng Yao, Min Zhang. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Jinsong Su, Deyi Xiong, Yang Liu 0005, Xianpei Han, Junfeng Yao, Min Zhang 0005 |
ACL (1) | 6 |
| 2015 | Graph-Based Collective Lexical Selection for Statistical Machine TranslationabstractLexical selection is of great importance to statistical machine translation. In this paper, we propose a graph-based frame-work for collective lexical selection. The framework is established on a translation graph that captures not only local associ-ations between source-side content words and their target translations but also target-side global dependencies in terms of relat-edness among target items. We also in-troduce a random walk style algorithm to collectively identify translations of source-side content words that are strongly related in translation graph. We validate the ef-fectiveness of our lexical selection frame-work on Chinese-English translation. Ex-periment results with large-scale training data show that our approach significantly improves lexical selection. 1 Jinsong Su, Deyi Xiong, Shujian Huang, Xianpei Han, Junfeng Yao |
EMNLP | 5 |
| 2015 | Bilingual Correspondence Recursive Autoencoder for Statistical Machine TranslationabstractLearning semantic representations and tree structures of bilingual phrases is beneficial for statistical machine translation.In this paper, we propose a new neural network model called Bilingual Correspondence Recursive Autoencoder (BCor-rRAE) to model bilingual phrases in translation.We incorporate word alignments into BCorrRAE to allow it freely access bilingual constraints at different levels.BCorrRAE minimizes a joint objective on the combination of a recursive autoencoder reconstruction error, a structural alignment consistency error and a crosslingual reconstruction error so as to not only generate alignment-consistent phrase structures, but also capture different levels of semantic relations within bilingual phrases.In order to examine the effectiveness of BCorrRAE, we incorporate both semantic and structural similarity features built on bilingual phrase representations and tree structures learned by BCorrRAE into a state-of-the-art SMT system.Experiments on NIST Chinese-English test sets show that our model achieves a substantial improvement of up to 1.55 BLEU points over the baseline. Jinsong Su, Deyi Xiong, Biao Zhang 0002, Yang Liu 0005, Junfeng Yao, Min Zhang 0005 |
EMNLP | 5 |
| 2015 | Shallow Convolutional Neural Network for Implicit Discourse Relation RecognitionabstractImplicit discourse relation recognition remains a serious challenge due to the absence of discourse connectives.In this paper, we propose a Shallow Convolutional Neural Network (SCNN) for implicit discourse relation recognition, which contains only one hidden layer but is effective in relation recognition.The shallow structure alleviates the overfitting problem, while the convolution and nonlinear operations help preserve the recognition and generalization ability of our model.Experiments on the benchmark data set show that our model achieves comparable and even better performance when comparing against current state-of-the-art systems. Biao Zhang 0002, Jinsong Su, Deyi Xiong, Yaojie Lu 0001, Hong Duan, Junfeng Yao |
EMNLP | 6 |
| 2015 | Discriminative Reordering Model Adaptation via Structural Learning
Biao Zhang 0002, Jinsong Su, Deyi Xiong, Hong Duan, Junfeng Yao |
IJCAI | 5 |
| 2015 | Realistic and stable animation of clothabstractAbstract Researchers have proposed various techniques for cloth simulation in the last decades. The crucial problem in interactive animation for cloth is how to speed up simulation with a stable system. In this paper, we describe a realistic and stable scheme for cloth simulation based on mass‐spring system. This scheme modifies semi‐implicit integration used in cloth simulation system with an efficient damping method. Semi‐implicit integration methods have been widely used in dynamic simulation because of acceptable speed and stability. However, internal damping forces are generated with respect to rotational rigid motions, which are only depending on relative velocities. Undesirable change in global movement of dynamic cloth could result in damping artifact. We replace the internal damping forces with an optimal damping method which is based on iterative spring damping but the limit can be computed directly. The method provides simple damping parameter estimation and guarantees conservation of global movement. As a result, complex clothes can be realistically and stably simulated in real time. Copyright © 2015 John Wiley & Sons, Ltd. Junfeng Yao, Lei Lan, Wenlin Lin, Yingying She |
Comput. Animat. Virtual Worlds | 1 |
| 2014 | Sparse Localized Decomposition of Deformation GradientsabstractAbstract Sparse localized decomposition is a useful technique to extract meaningful deformation components out of a training set of mesh data. However, existing methods cannot capture large rotational motion in the given mesh dataset. In this paper we present a new decomposition technique based on deformation gradients. Given a mesh dataset, the deformation gradient field is extracted, and decomposed into two groups: rotation field and stretching field, through polar decomposition. These two groups of deformation information are further processed through the sparse localized decomposition into the desired components. These sparse localized components can be linearly combined to form a meaningful deformation gradient field, and can be used to reconstruct the mesh through a least squares optimization step. Our experiments show that the proposed method addresses the rotation problem associated with traditional deformation decomposition techniques, making it suitable to handle not only stretched deformations, but also articulated motions that involve large rotations. Junfeng Yao, Zichun Zhong, Yang Liu 0013, Xiaohu Guo |
Comput. Graph. Forum | 2 |
| 2013 | Creating a large area of trees based on FOREST PROabstractForest Pro makes you can create large area trees, bushes and other models in a short time. If you want to create 50000 trees in a scene, with each tree has more than 10 thousands multilateral type, even if the computer can drag, it may get stuck in the rendering time. Now create tens of thousands of trees through Forest Pro is easy, but also it won't get stuck in the rendering process, this is a big edge for the animators who often make outdoor landscape. Made by IToo company. As is shown in figure 1. Jincan Lin, Yunning Huang, Junfeng Yao |
I3D | 3 |
| 2011 | The research and realization of parameterized three-dimensional plant simulation based on fractal theoryabstractThe simulation of plant morphology has become a research hotspot in computer graphic. Two methods are commonly used to simulate plant morphology: L-system and IFS (Iterated Function System). However, they are limited to the construction methods of fractal plants. Moreover, the degree of realism in three-dimensional simulation is still insufficient, and the morphology of different plants has small differences. Junfeng Yao, Fengchun Lin, Binxing Wang |
SI3D | 1 |
| 2010 | Research and application of 3D engine in the simulation of natural sceneabstractNo abstract available. Junfeng Yao, Andy Ju An Wang |
SI3D | 1 |
| 2010 | A new method of information decision-making based on D-S evidence theoryabstractD-S evidence theory is a method broadly applied in fusion for decision-making. However, this theory has some shortcomings in the formula of evidence combination with the exception that evidence of fully conflict can not be combined, then the probability validity is difficult to determine and sometimes the composed evidence is different from people's subjective judgments or some other issues. These confine the application of evidence to some extent. Some of them have the dubious credibility which affects the fusion result when Multi evidence are combined together. In order to expand the application of the formula of this theory and enhance the reliability of the fusion results, a new combination formula is introduced in this paper, which is also compared with other formulas in other literatures and finally the improved reliability of this combination formula is verified. At last, through data-mining of the decision-making information on a number of isolated points, a new method using combined evidence to make decisions is described. It's proven from the experimental results that the new combination method not only works well and effectively in the evidence of a high level of conflict but also is applicable to fusion for decision-making. Junfeng Yao, Chengpeng Wu, Xiaobiao Xie, Guoli Ji, Prabir Bhattacharya |
SMC | 1 |
| 2008 | Design and Implementation of the Container Terminal Operating System Based on Service-Oriented Architecture (SOA)abstractDue to the impact on IT architecture, the limitation of information planning, as well as the problems left over by history, etc., traditional terminal operating system are often unable to make rapid response to the multi-branch expansion, and unable to adapt and digest the impact on terminal operations made by the frequent changes in business processes. Firstly, on the basis of analyzing the container terminal business processes in detail, this paper sorted out the sequence of key processes. Then, on the basis of SOA model, combined with layered theory, through the design and combinational calling of Web services, the paper implemented a set of pivotal business processes. The characteristic of this paper is that the SOA theory is applied to terminal processes of container port industry. Integrated the general principles of SOA with the characteristics of container port industry, the paper built a more viable IT architecture, completed the Web services design of pivotal container terminal business processes, and realized the design practice of container terminal operating system. Junfeng Yao |
CW | 2 |
| 2008 | The Technical Research and System Realization of 3D Dice GameabstractDice is familiar to everyone, but there is few applied 3D dice game. This paper introduces the key technology of 3D dice game developed by C++ and OpenGL. Itpsilas possible to run the game on embedded in future as abundant mathematic knowledge is performed to the 3D dice game that collision effect would be more realistic and dices would run with high speed.This paper discusses how to conquer dice-circle collision and dice-dice collision, how to draw 3d dice with round corners and a simple cylinder in which dices will bounce . Junfeng Yao, Yunhui Huang |
CW | 1 |
| 2008 | Research on Method of 3D Reconstruction of Ancient Architecture (Nanputuo Temple)abstractIn this paper, regarding the Great Majesty Hall of the NanPutuo Temple as a virtual modeling object, the research on the new technology of 3D reconstruction of the ancient architecture has been done, in which 3DS MAX and MultiGen Creator were applied. The conflict was solved between precision and the amount of data existing in the process of 3D modeling virtual ancient architecture. The practical project has proved that the ancient architecture model made in this way has gotten living effect in roaming system, at the mean time it satisfies the data demand of real-time rendering. Junfeng Yao, Fei She |
CW | 1 |
| 2008 | Study on Bezier Curve Applications of Pelvis Trajectory of Virtual Human WalkingabstractAs the root joint of the whole lower limbs, pelvis trajectory decides the motion trajectory of the virtual human walking. Most of the current methods of simulating pelvis trajectory can work reasonably well on even terrain, but on the uneven terrain, these methods appear insufficient. This paper uses Bezier curve to simulate the pelvis trajectory. We attain the motion trajectory by using the position points in the middle of leg supporting duration and in the middle of double support phase as control points, then connecting the curve. This approach not only work reasonably well in modeling human normal walking on even terrain, but also appear sufficient to generate human walking on uneven terrain. Junfeng Yao, Hanhui Zhang, Qingqing Cheng |
CW | 1 |