Zhenghan Chen

dblp:192/7651 · DBLP profile ↗
← Back
28ranked-venue papers
1as first author
27since 2021 · last 2026
0000-0002-1841-539XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 15 since 2021Artificial intelligence and machine learning · 13 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A sign language to SQL query translation system for enhancing database accessibility
Guocang Yang, Dawei Yuan, Tao Zhang 0001, Zhenghan Chen
Autom. Softw. Eng.4
2025 Reducing Divergence in Batch Normalization for Domain Adaptation
abstract
The widespread adoption of Batch Normalization (BN) in contemporary deep neural architectures has demonstrated significant efficacy, particularly in the domain of Unsupervised Domain Adaptation (UDA) for cross-domain applications. Notwithstanding its success, extant BN variants often conflate source and target domain information within identical channels, potentially compromising transferability due to inter-domain feature misalignment. To address this limitation, we introduce Refined Batch Normalization (RBN), a novel normalization paradigm that leverages estimated shift to quantify discrepancies between estimated population statistics and their expected values. Our pivotal observation reveals that estimated shift can accumulate through BN stacking within the network, potentially degrading target domain performance. We elucidate how RBN mitigates this accumulation, thereby enhancing overall system efficacy. The practical implementation of this technique is realized through the RBNBlock, which supplants conventional BN with RBN in the bottleneck architecture of residual networks. Extensive empirical evaluation across diverse cross-domain benchmarks corroborates the superiority of RBN in augmenting inter-domain transferability. This perspective transcends immediate performance metrics, offering a foundational lens through which subsequent research can more deeply understand and refine the interplay between normalization strategies and domain adaptation.
Ellen Yi-Ge, Mingjing Wu, Zhenghan Chen
AAAI3
2025 Image-to-video Adaptation with Outlier Modeling and Robust Self-learning
abstract
The image-to-video adaptation task seeks to effectively harness both labeled images and unlabeled videos for achieving effective video recognition. The modality gap of the image and video modalities and the domain discrepancy across the two domains are the two essential challenges in this task. Existing methods reduce the domain discrepancy via close-set domain adaptation techniques, resulting in inaccurate domain alignment as there exist outlier target frames. To tackle this issue, we extend the vanilla classifier with outlier classes, where each outlier class responsible for capturing outlier frames for a specific class via batch nuclear norm maximization loss. We further propose a new loss by treating the source images apart from class c as instances from outlier class specific for c. As for the modality gap, existing methods usually utilize the pseudo labels obtained from an image-level adapted model to learn a video-level model. Rare efforts are dedicated to handling the noise in pseudo labels. We proposed a new metric based on label propagation consistency to select samples for training a better video-level model. Experiments on 3 benchmarks validating the effectiveness of our method.
Junbao Zhuo, Shuhui Wang, Zhenghan Chen, Li Shen 0005, Qingming Huang, Huimin Ma 0001
AAAI3
2025 Towards a 3D Transfer-Based Black-Box Attack via Critical Feature Guidance
abstract
Deep neural networks for 3D point clouds have been demonstrated to be vulnerable to adversarial examples. Previous 3D adversarial attack methods often exploit certain information about the target models, such as model parameters or outputs, to generate adversarial point clouds. However, in realistic scenarios, it is challenging to obtain any information about the target models under conditions of absolute security. Therefore, we focus on transfer-based attacks, where generating adversarial point clouds does not require any information about the target models. Based on our observation that the critical features used for point cloud classification are consistent across different DNN architectures, we propose CFG, a novel transfer-based black-box attack method that improves the transferability of adversarial point clouds via the proposed Critical Feature Guidance. Specifically, our method regularizes the search of adversarial point clouds by computing the importance of the extracted features, prioritizing the corruption of critical features that are likely to be adopted by diverse architectures. Further, we explicitly constrain the maximum deviation extent of the generated adversarial point clouds in the loss function to ensure their imperceptibility. Extensive experiments conducted on the ModelNet40 and ScanObjectNN benchmark datasets demonstrate that the proposed CFG outperforms the state-of-the-art attack methods by a large margin.
Shuchao Pang, Zhenghan Chen, Siyuan Liang 0004, Anan Du, Yongbin Zhou
ICCV2
2025 Natural Humanoid Robot Locomotion with Generative Motion Prior
abstract
Natural and lifelike locomotion remains a fundamental challenge for humanoid robots to interact with human society. However, previous methods either neglect motion naturalness or rely on unstable and ambiguous style rewards. In this paper, we propose a novel Generative Motion Prior (GMP) that provides fine-grained motion-level supervision for the task of natural humanoid robot locomotion. To leverage natural human motions, we first employ whole-body motion retargeting to effectively transfer them to the robot. Subsequently, we train a generative model offline to predict future natural reference motions for the robot based on a conditional variational auto-encoder. During policy training, the generative motion prior serves as a frozen online motion generator, delivering precise and comprehensive supervision at the trajectory level, including joint angles and keypoint positions. The generative motion prior significantly enhances training stability and improves interpretability by offering detailed and dense guidance throughout the learning process. Experimental results in both simulation and real-world environments demonstrate that our method achieves superior motion naturalness compared to existing approaches. Project page can be found at https://sites.google.com/view/humanoid-gmp
Zhenghan Chen, Yue Wang 0020, Rong Xiong
IROS3
2025 CbDA: Contrastive-Based Data Augmentation for Domain Generalization
abstract
In the realm of domain generalization (DG), domain adversarial training is a popular method for achieving invariant representations and is often applied to various tasks in this field. Notably, recent developments in supervised learning, particularly in classification, have shown that methods converging toward smoother optima yield better generalization. This research delves into the impact of contrastive-based data augmentation on DG, focusing on leveraging category-specific distribution statistics. We introduce an innovative contrastive loss at the sample level, tailored to align samplewise representations with semantic distributions across domains. This involves encouraging representations within the same category to form clusters while ensuring those from different categories remain distinct, thus enhancing the model's classification strength. Additionally, we establish an upper limit for this loss function. This approach efficiently handles an infinite array of both similar and dissimilar sample pairs. Our methodology significantly surpasses the baseline model, a fact underscored by comprehensive empirical evaluations on challenging benchmarks such as Digits-DG, PACS, Office–Home, and DomainNet.
Xiaoxuan Liang 0002, Zhenghan Chen
IEEE Trans. Comput. Soc. Syst.4
2024 Sparse Enhanced Network: An Adversarial Generation Method for Robust Augmentation in Sequential Recommendation
abstract
Sequential Recommendation plays a significant role in daily recommendation systems, such as e-commerce platforms like Amazon and Taobao. However, even with the advent of large models, these platforms often face sparse issues in the historical browsing records of individual users due to new users joining or the introduction of new products. As a result, existing sequence recommendation algorithms may not perform well. To address this, sequence-based data augmentation methods have garnered attention. Existing sequence enhancement methods typically rely on augmenting existing data, employing techniques like cropping, masking prediction, random reordering, and random replacement of the original sequence. While these methods have shown improvements, they often overlook the exploration of the deep embedding space of the sequence. To tackle these challenges, we propose a Sparse Enhanced Network (SparseEnNet), which is a robust adversarial generation method. SparseEnNet aims to fully explore the hidden space in sequence recommendation, generating more robust enhanced items. Additionally, we adopt an adversarial generation method, allowing the model to differentiate between data augmentation categories and achieve better prediction performance for the next item in the sequence. Experiments have demonstrated that our method achieves a remarkable 4-14% improvement over existing methods when evaluated on the real-world datasets. (https://github.com/junyachen/SparseEnNet)
Junyang Chen 0001, Guoxuan Zou, Pan Zhou 0001, Yirui Wu, Zhenghan Chen, Houcheng Su, Huan Wang 0005, Zhiguo Gong
AAAI5
2024 UniADS: Universal Architecture-Distiller Search for Distillation Gap
abstract
In this paper, we present UniADS, the first Universal Architecture-Distiller Search framework for co-optimizing student architecture and distillation policies. Teacher-student distillation gap limits the distillation gains. Previous approaches seek to discover the ideal student architecture while ignoring distillation settings. In UniADS, we construct a comprehensive search space encompassing an architectural search for student models, knowledge transformations in distillation strategies, distance functions, loss weights, and other vital settings. To efficiently explore the search space, we utilize the NSGA-II genetic algorithm for better crossover and mutation configurations and employ the Successive Halving algorithm for search space pruning, resulting in improved search efficiency and promising results. Extensive experiments are performed on different teacher-student pairs using CIFAR-100 and ImageNet datasets. The experimental results consistently demonstrate the superiority of our method over existing approaches. Furthermore, we provide a detailed analysis of the search results, examining the impact of each variable and extracting valuable insights and practical guidance for distillation design and implementation.
Zhenghan Chen, Yihang Rao, Lujun Li 0001, Shuchao Pang
AAAI2
2024 GSENet: Global Semantic Enhancement Network for Lane Detection
abstract
Lane detection is the cornerstone of autonomous driving. Although existing methods have achieved promising results, there are still limitations in addressing challenging scenarios such as abnormal weather, occlusion, and curves. These scenarios with low visibility usually require to rely on the broad information of the entire scene provided by global semantics and local texture information to predict the precise position and shape of the lane lines. In this paper, we propose a Global Semantic Enhancement Network for lane detection, which involves a complete set of systems for feature extraction and global features transmission. Traditional methods for global feature extraction usually require deep convolution layer stacks. However, this approach of obtaining global features solely through a larger receptive field not only fails to capture precise global features but also leads to an overly deep model, which results in slow inference speed. To address these challenges, we propose a novel operation called the Global feature Extraction Module (GEM). Additionally, we introduce the Top Layer Auxiliary Module (TLAM) as a channel for feature distillation, which facilitates a bottom-up transmission of global features. Furthermore, we introduce two novel loss functions: the Angle Loss, which account for the angle between predicted and ground truth lanes, and the Generalized Line IoU Loss function that considers the scenarios where significant deviations occur between the prediction of lanes and ground truth in some harsh conditions. The experimental results reveal that the proposed method exhibits remarkable superiority over the current state-of-the-art techniques for lane detection.Our codes are available at:https://github.com/crystal250/GSENet.
Junhao Su, Zhenghan Chen, Chenghao He, Dongzhi Guan, Changpeng Cai, Tongxi Zhou, Jiashen Wei, Wenhua Tian, Zhihuai Xie
AAAI2
2024 Sharpness-Aware Model-Agnostic Long-Tailed Domain Generalization
abstract
Domain Generalization (DG) aims to improve the generalization ability of models trained on a specific group of source domains, enabling them to perform well on new, unseen target domains. Recent studies have shown that methods that converge to smooth optima can enhance the generalization performance of supervised learning tasks such as classification. In this study, we examine the impact of smoothness-enhancing formulations on domain adversarial training, which combines task loss and adversarial loss objectives. Our approach leverages the fact that converging to a smooth minimum with respect to task loss can stabilize the task loss and lead to better performance on unseen domains. Furthermore, we recognize that the distribution of objects in the real world often follows a long-tailed class distribution, resulting in a mismatch between machine learning models and our expectations of their performance on all classes of datasets with long-tailed class distributions. To address this issue, we consider the domain generalization problem from the perspective of the long-tail distribution and propose using the maximum square loss to balance different classes which can improve model generalizability. Our method's effectiveness is demonstrated through comparisons with state-of-the-art methods on various domain generalization datasets. Code: https://github.com/bamboosir920/SAMALTDG.
Houcheng Su, Weihao Luo, Daixian Liu, Mengzhu Wang, Junyang Chen 0001, Cong Wang 0018, Zhenghan Chen
AAAI8
2024 Dynamic Spiking Graph Neural Networks
abstract
The integration of Spiking Neural Networks (SNNs) and Graph Neural Networks (GNNs) is gradually attracting attention due to the low power consumption and high efficiency in processing the non-Euclidean data represented by graphs. However, as a common problem, dynamic graph representation learning faces challenges such as high complexity and large memory overheads. Current work often uses SNNs instead of Recurrent Neural Networks (RNNs) by using binary features instead of continuous ones for efficient training, which overlooks graph structure information and leads to the loss of details during propagation. Additionally, optimizing dynamic spiking models typically requires the propagation of information across time steps, which increases memory requirements. To address these challenges, we present a framework named Dynamic Spiking Graph Neural Networks (Dy-SIGN). To mitigate the information loss problem, Dy-SIGN propagates early-layer information directly to the last layer for information compensation. To accommodate the memory requirements, we apply the implicit differentiation on the equilibrium state, which does not rely on the exact reverse of the forward computation. While traditional implicit differentiation methods are usually used for static situations, Dy-SIGN extends it to the dynamic graph setting. Extensive experiments on three large-scale real-world dynamic graph datasets validate the effectiveness of Dy-SIGN on dynamic node classification tasks with lower computational costs.
Mengzhu Wang, Zhenghan Chen, Giulia De Masi, Huan Xiong, Bin Gu 0001
AAAI3
2024 Multi-Level Adaptive Graph Networks for Brain MRI Classification in Neurodegenerative Disorders
abstract
We present an innovative methodology for the classification of Presenile dementia (Pd) using the Hierarchical Adaptive Multi-Domain Analysis Network (HAMDA-Net). This architecture effectively combines sophisticated mathematical theories with cutting-edge domain adaptation techniques, transforming the analysis of functional magnetic resonance imaging (fMRI) data. Within the scope of domain adaptation, we introduce a novel intra-group regularization component that significantly enhances the model’s capacity to generalize across diverse patient populations and varying imaging protocols. This advancement addresses a critical challenge in medical imaging analysis, where inter-site variability often impedes the broad applicability of machine learning models. We provide theoretical assurances regarding the convergence properties of our domain adaptation strategy, demonstrating its effectiveness in minimizing dataset bias while preserving the discriminative features essential for accurate Pd classification. In conclusion, HAMDA-Net represents a transformative advancement in the application of deep learning to neuroimaging analysis. By seamlessly integrating advanced mathematical frameworks with innovative domain adaptation methodologies, our work not only extends the frontiers of Pd classification but also establishes a robust framework adaptable to a wide range of neurological disorders. The theoretical foundations and empirical validations presented herein pave the way for the development of more accurate, interpretable, and generalizable models in medical imaging analysis, with profound implications for personalized healthcare and neuroscientific research.
Jingyuan Dai, Zhenghan Chen
BIBM3
2024 Cross-Site Adaptation for COVID-19 Diagnosis Using Graph-Based Neural Networks
abstract
The SARS-CoV-2 outbreak has precipitated an unparalleled global health crisis, affecting numerous nations worldwide. As the incidence of new infections persists in rising, there is an urgent necessity for automated systems capable of detecting SARS-CoV-2 through computed tomography (CT) imaging. Such systems possess considerable potential to support clinical diagnostics and alleviate the substantial burden linked to manual image analysis. Enhancing the datasets utilized for developing machine learning models necessitates the inclusion of cases from diverse medical systems. This approach is essential for constructing models that are both robust and generalizable. In this investigation, we present an innovative methodology named Domain Adaptive Graph Alignment with Neural Network (DAGANN) for the accurate identification of SARS-CoV-2. DAGANN is engineered to effectively learn from heterogeneous datasets characterized by distributional variations, employing the resilient architecture of COVID-Net to boost both prediction accuracy and learning efficiency. The DAGANN model offers several significant contributions. Primarily, to our knowledge, this study is the first to incorporate graph-based information within a deep learning framework for SARS-CoV-2 detection. Additionally, the model integrates three distinct alignment mechanisms designed to learn domain-invariant and semantic representations, thereby reducing domain discrepancies and facilitating domain adaptation. We validate our approach using two publicly available, large-scale SARS-CoV-2 diagnostic datasets composed of CT images. Comprehensive experimental results demonstrate that our method consistently achieves superior performance across both datasets. Furthermore, our approach surpasses current leading multi-site learning techniques, underscoring its efficacy and superiority.
Mingshuo Wang, Yiding Cui, Zhenghan Chen
BIBM3
2024 Cross-Domain Transfer in Residual Networks for Clinical Image Partitioning
abstract
The accurate recognition and comprehensive understanding of medical images depicting human tissue represent a central focus in computer vision research. Many tasks within medical imaging rely on deep neural networks, particularly those with U-shaped architectures and skip connections. The advancement of computer vision technologies demands the application of convolutional neural networks (CNNs). Despite progress, two major challenges remain in medical image processing: (1) developing a model framework with low computational complexity that allows for efficient inference without compromising accuracy, and (2) designing a model with strong generalization capabilities across various datasets derived from patients with differing pathologies, thereby mitigating domain shift challenges. In response to the first issue, we propose a novel unsupervised domain adaptation method utilizing Interoperable Batch Normalization (IBN) to integrate multiple channels within deep neural networks, enhancing adversarial domain adaptation. Our experimental evaluation on the Hubmap and Synapse multiorgan segmentation datasets reveals that the proposed RRUNet model achieves superior performance compared to existing methods, setting a new standard in the domain.
Mingshuo Wang, Keyan Jin, Wenzhuo Bao, Zhenghan Chen
BIBM4
2024 Adaptive Learning for Medical Image Classification across Domains
abstract
In the rapidly advancing area of semi-supervised learning (SSL), especially concerning large-scale computer vision models and medical imaging, pseudo-labeling has become a central strategy. Nevertheless, this technique faces significant hurdles, including the possibility of integrating erroneously pseudo-labeled data due to improper threshold settings and the underutilization of data with low-confidence pseudo-labels. To overcome these challenges, we introduce a novel SSL framework for training large models, termed PE-DAT, which fuses Pseudo-loss Estimation with Domain-Adaptive Training. This integration substantially enhances the accuracy of medical image classification tasks, addressing both multiclass and multi-label problems. A pivotal aspect of our approach is a cutting-edge data selection method based on the concept that genuine data exhibit lower pseudo-loss values compared to their noisy counterparts. The PE-DAT framework’s effectiveness is demonstrated through its superior performance relative to existing benchmarks, validated by extensive testing on a wide array of medical and natural image datasets.
Mingshuo Wang, Yong Zhuang, Zhenghan Chen
BIBM4
2024 Mitigating Feature Homogenization in Deep Graph Architectures for Clinical Data Representation
abstract
This research redefines International Classification of Diseases (ICD) coding as a sophisticated multi-label prediction problem, requiring the assignment of multiple codes to detailed discharge summaries. Current automated ICD coding techniques face challenges in effectively classifying medical diagnostic texts that involve complex and sparse label distributions, especially when model parameters are adjusted using traditional backpropagation methods. We present LGG-NRGrand, a novel adversarial framework that approaches ICD coding through the generation of labeled graphs. A significant challenge in this field is the widespread issue of Over-Smoothing in deep graph neural networks, which results in uniform or indistinct node representations. Our model is designed to improve the capacity for learning heterogeneous graph representations within a layered network architecture. Specifically, we introduce NRGrand, a single-relational deep graph neural network structure that mitigates the Over-Smoothing problem while capturing more detailed graph features during the representation learning phase. The LGG-NRGrand model is trained using an adversarial reinforcement framework, employing an adversarial domain adaptation technique. Experimental results indicate that LGG-NRGrand surpasses current methods on key evaluation metrics, including micro-F1, micro-AUC, and P@K.
Suyang Xi, Bolin Yang, Zhenghan Chen
BIBM3
2024 Federated Learning on Distributed Graphs Considering Multiple Heterogeneities
abstract
Federated graph learning (FGL) collaboratively learns a global graph neural network with distributed graphs, where a significant challenge is addressing non-IID issues. Existing work has not fully explored and utilized the intrinsic features of graphs, resulting in their inability to effectively solve non-IID issues. To tackle this challenge, we investigate for the first time the various heterogeneity that causes non-IID issues in FGL and how they can be utilized to alleviate the issues, including the heterogeneity of nodes and structures as basic components of the graph, as well as the resulting heterogeneity in the representations of the graph. Furthermore, we propose ProtoFGL to address these issues. ProtoFGL first extracts heterogeneous features of nodes and structures from local data and incorporates them into prototypes, which are then used as graph representations for collaborative training. Experimental results show that ProtoFGL outperforms state-of-the-art methods in node classification tasks in accuracy and F1 score.
Yedi Ma, Hongyan Gu, Zhenghan Chen, Xinli Huang
ICASSP5
2024 Domain Adaptive Graph Classification
abstract
Despite the remarkable accomplishments of graph neural networks (GNNs), they typically rely on task-specific labels, posing potential challenges in terms of their acquisition. Existing work have been made to address this issue through the lens of unsupervised domain adaptation, wherein labeled source graphs are utilized to enhance the learning process for target data. However, the simultaneous exploration of graph topology and reduction of domain disparities remains a substantial hurdle. In this paper, we introduce the Dual Adversarial Graph Representation Learning (DAGRL), which explore the graph topology from dual branches and mitigate domain discrepancies via dual adversarial learning. Our method encompasses a dual-pronged structure, consisting of a graph convolutional network branch and a graph kernel branch, which enables us to capture graph semantics from both implicit and explicit perspectives. Moreover, our approach incorporates adaptive perturbations into the dual branches, which align the source and target distribution to address domain discrepancies. Extensive experiments on a wild range graph classification datasets demonstrate the effectiveness of our proposed method.
Siyang Luo, Zhenghan Chen, Xiaoxuan Liang 0002
ICASSP3
2024 Semanticmapper: Region-Specific Domain Adaptation for 3D Shapes Through Lexical Delineation
abstract
In recent advancements within the domain of three-dimensional semantic mapping, a novel framework termed SemanticMapper has emerged, heralding a paradigm shift in the arena of mesh localization through textual inputs. The core innovation of SemanticMapper lies in its adeptness at navigating and accurately projecting semantic regions onto meshes, particularly excelling in "out-of-domain" scenarios. This capability is exemplified in its seamless overlay of attire onto unclad three-dimensional animal forms. The operational backbone of SemanticMapper is a neural field, ingeniously designed to interpret and apply context to textual descriptions. This is synergistically coupled with a sophisticated probability-weighted blending algorithm, which ensures precise coloration of targeted mesh regions. A pivotal advantage of SemanticMapper is its autonomy from the requisites of three-dimensional data sets or manual annotations, a feature made possible by incorporating a pre-trained CLIP2 encoder. A cornerstone of this research is the integration of Combined Adversarial Domain Adaptation (CADA), a mechanism that significantly enhances the domain adaptability of SemanticMapper. This advancement not only extends the system’s efficacy but also broadens its applicability across a diverse array of input shapes, marking a substantial leap in the field of semantic mesh localization.
Tianci Xie, Siyang Luo, Zhenghan Chen, Xiaoxuan Liang 0002
ICASSP3
2024 DREAM: Dual Structured Exploration with Mixup for Open-set Graph Domain Adaption
abstract
Recently, numerous graph neural network methods have been developed to tackle domain shifts in graph data. However, these methods presuppose that unlabeled target graphs belong to categories previously seen in the source domain. This assumption could not hold true for in-the-wild target graphs. In this paper, we delve deeper to explore a more realistic problem open-set graph domain adaptation. Our objective is to not only identify target graphs from new categories but also accurately classify remaining target graphs into their respective categories under domain shift and label scarcity. To solve this challenging problem, we introduce a new method named Dual Structured Exploration with Mixup (DREAM). DREAM incorporates a graph-level representation learning branch as well as a subgraph-enhanced branch, which jointly explores graph topological structures from both global and local viewpoints. To maximize the use of unlabeled target graphs, we train these two branches simultaneously using posterior regularization to enhance their inter-module consistency. To accommodate the open-set setting, we amalgamate dissimilar samples to generate virtual unknown samples belonging to novel classes. Moreover, to alleviate domain shift, we establish a k nearest neighbor-based graph-of-graphs and blend multiple neighbors of each sample to produce cross-domain virtual samples for inter-domain consistency learning. Extensive experiments validate the effectiveness of the proposed DREAM in comparison to various state-of-the-art approaches in different settings.
Mengzhu Wang, Zhenghan Chen, Li Shen 0008, Huan Xiong, Bin Gu 0001, Xiao Luo 0001
ICLR3
2024 A Fast Motion and Foothold Planning Framework for Legged Robots on Discrete Terrain
abstract
Legged robot proved their capability to cross complex terrain in recent research, yet the autonomy of robots on discrete terrain still needs to be enhanced since it requires a full stack framework. This paper introduces a real-time motion and foothold planning framework tailored for legged robots navigating uneven terrains, such as stepping stones. Our approach addresses the critical challenges of determining feasible global paths and local footholds to enhance autonomous mobility across complex landscapes. By using a sampling-based global path planner integrated with terrain segmentation and the robot’s kinematic model, our framework swiftly generates viable navigation paths. Concurrently, it utilizes a Mixed Integer Programming (MIP) methodology for real-time foothold optimization, ensuring the robot’s stability and safety through dynamic terrain interaction. Finally, an execution layer including Model Predictive Control (MPC) and Whole-Body Control (WBC) generates the robots’ motion. Simulation and real-world experiments demonstrate that our framework improves legged robots’ adaptability on discrete terrains.
Jiyu Yu, Dongqi Wang 0002, Zhenghan Chen, Ci Chen 0004, Shuangpeng Wu, Yue Wang 0020, Rong Xiong
IROS3
2024 Improving diversity and discriminability based implicit contrastive learning for unsupervised domain adaptation
Chuanqi Shi, WenZe Fan, Zhenghan Chen
Appl. Intell.4
2024 Multidocument Aspect Classification for Aspect-Based Abstractive Summarization
abstract
Multidocument aspect-based summarization (AspSumm) aims to generate focused summaries based on the target aspects from a cluster of relevant documents. Generating such summaries can better satisfy readers’ specific points of interest, as readers may have different concerns about the same articles. However, previous methods usually generate aspect-based summaries based on the given aspects without using the relationship among aspects to assist in the summarization. In this work, we propose a two-stage general framework for multidocument AspSumm. The model first discovers the latent relationship among aspects and then uses relevant sentences selected by aspect discovery to generate abstractive summaries. We exploit latent dependencies among aspects using a tag mask training (TMT) strategy, which increases the interpretability of the model. In addition to improvements in summarization over aspect-based strong baselines, experimental results show that our proposed model can accurately discover multidomain aspects on the WikiAsp dataset.
Ye Wang 0023, Mengzhu Wang, Zhenghan Chen, Zhiping Cai, Junyang Chen 0001, Victor C. M. Leung
IEEE Trans. Comput. Soc. Syst.4
2023 HAG: Hierarchical Attention with Graph Network for Dialogue Act Classification in Conversation
abstract
The prediction of dialogue acts (DA) labels on utterance-level in conversations can be treated as a sequence labeling problem, which requires context- and speaker-aware semantic comprehension, especially for Japanese. In this study, we pro-posed a hierarchical attention with the graph neural network (HAG) to consider the contextual interconnections as well as the semantics carried by the sentence itself. Concretely, the model use long-short term memory networks (LSTMs) to perform a context-aware encoding within a dialogue window. Then, we construct the context graph by aggregating the neighboring utterances. Subsequently, a speaker feature transformation is executed with a graph attention network (GAT) to calculate the interconnections, while a context-level feature selection is performed with a gated graph convolutional network (GatedGCN) to select the salient utterances that contribute to the DA classification. Finally, we merge the representations of different levels and conduct a classification with two dense layers. We evaluate the proposed model on Japanese dialogue act dataset (JPS-DA). The experimental results show that our method outperforms the baselines.
Changzeng Fu, Zhenghan Chen, Bowen Wu 0002, Carlos Toshinori Ishi, Hiroshi Ishiguro
ICASSP2
2023 LGFat-RGCN: Faster Attention with Heterogeneous RGCN for Medical ICD Coding Generation
abstract
With the increasing volume of healthcare data, automated International Classification of Diseases (ICD) has become increasingly relevant and is frequently regarded as a medical multi-label prediction problem. Current methods struggle to accurately classify medical diagnosis texts that represent deep and sparse categories. Unlike these works that model the label with code hierarchy or description for label prediction, we argue that the label generation with structural information can provide more comprehensive knowledge based on the observation that label synonyms and parent-child relationships in vary from their context in clinical contexts. In this study, we introduce \tool, a heterogeneous graph model with improved attention for automated ICD coding. Notably, our approach represents the model to consider this task as a labelled graph generation problem. Our enhanced attention mechanism boosts the model's capacity to learn from multi-relational heterogeneous graph representations. Additionally, we propose a discriminator for labelled graphs (LG) that computes the reward for each ICD code in the labelled graph generator. Our experimental findings demonstrate that our proposed model significantly outperforms all existing strong baseline methods and attains the best performance on three benchmark datasets.
Zhenghan Chen, Changzeng Fu, Ruoxue Wu, Ye Wang 0023, Xunzhu Tang, Xiaoxuan Liang 0002
ACM Multimedia1
2023 TE-KWS: Text-Informed Speech Enhancement for Noise-Robust Keyword Spotting
abstract
Keyword spotting (KWS) presents a formidable challenge, particularly in high-noise environments. Traditional denoising algorithms that rely solely on speech have difficulty recovering speech that has been severely corrupted by noise. In this investigation, we develop an adaptive text-informed denoising model to bolster reliable keyword identification in the presence of considerable noise degradation. The whole proposed TE-KWS incorporates a tripartite branch structure, where the speech branch (SB) takes noisy speech as input which provides the raw speech information, the alignment branch (AB) accommodates aligned text input which facilitates accurate restoration of the corresponding speech when text with alignment is preserved, and the text branch (TB) handles unaligned text which prompts the model to autonomously learn the alignment between speech and text. To make the proposed denoising model more beneficial for KWS, following the training of the whole model,the alignment branch (AB) is frozen, and the model is fine-tuned by leveraging its speech restoration and forced alignment capabilities. Subsequently, the input for the text branch (TB) is supplanted with designated keywords, and a heavier denoising penalty is applied on the keywords period, thereby explicitly intensifying the speech restoration ability of the model for keywords. Finally, the Combined Adversarial Domain Adaptation (CADA) is implemented to enhance the robustness of KWS with regard to data pre-and post-speech enhancement (SE). Experimental results indicate that our approach not only markedly ameliorates highly corrupted speech, achieving SOTA performance for marginally corrupted speech, but also bolsters the efficacy and generalizability of prevailing mainstream KWS models.
Dong Liu 0037, Qirong Mao, Lijian Gao, Qinghua Ren, Zhenghan Chen, Ming Dong 0001
ACM Multimedia5
2023 A Closer Look at Classifier in Adversarial Domain Generalization
abstract
The task of domain generalization is to learn a classification model from multiple source domains and generalize it to unknown target domains. The key to domain generalization is learning discriminative domain-invariant features. Invariant representations are achieved using adversarial domain generalization as one of the primary techniques. For example, generative adversarial networks have been widely used, but suffer from the problem of low intra-class diversity, which can lead to poor generalization ability. To address this issue, we propose a new method called auxiliary classifier in adversarial domain generalization (CloCls). CloCls improve the diversity of the source domain by introducing auxiliary classifier. Combining typical task-related losses, e.g., cross-entropy loss for classification and adversarial loss for domain discrimination, our overall goal is to guarantee the learning of condition-invariant features for all source domains while increasing the diversity of source domains. Further, inspired by smoothing optima have improved generalization for supervised learning tasks like classification. We leverage that converging to a smooth minima with respect task loss stabilizes the adversarial training leading to better performance on unseen target domain which can effectively enhances the performance of domain adversarial methods. We have conducted extensive image classification experiments on benchmark datasets in domain generalization, and our model exhibits sufficient generalization ability and outperforms state-of-the-art DG methods.
Ye Wang 0023, Junyang Chen 0001, Mengzhu Wang, Hao Li 0058, Wei Wang 0335, Houcheng Su, Zhihui Lai 0001, Wei Wang 0077, Zhenghan Chen
ACM Multimedia9
2017 Cooperative particle swarm optimization using MapReduce
Yang Wang 0075, Yangyang Li 0001, Zhenghan Chen, Yu Xue 0003
Soft Comput.3