EDBT 2026 Demo / reviewers in the wild / expert
Wei Peng 0011
dblp:16/5560-11
· DBLP profile ↗
35ranked-venue papers
2as first author
30since 2021 · last 2025
0000-0002-0868-0974ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 1 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 12 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering VectorsabstractLarge language models (LLMs) have achieved remarkable performance across many tasks, yet aligning them with desired behaviors remains challenging. Activation intervention has emerged as an effective and economical method to modify the behavior of LLMs. Despite considerable interest in this area, current intervention methods exclusively employ a fixed steering vector to modify model activations, lacking adaptability to diverse input semantics. To address this limitation, we propose Semantics-Adaptive Dynamic Intervention (SADI), a novel method that constructs a dynamic steering vector to intervene model activations at inference time. More specifically, SADI utilizes activation differences in contrastive pairs to precisely identify critical elements of an LLM (i.e., attention heads, hidden states, and neurons) for targeted intervention. During inference, SADI dynamically steers model behavior by scaling element-wise activations based on the directions of input semantics. Experimental results show that SADI outperforms established baselines by substantial margins, improving task performance without training. SADI's cost-effectiveness and generalizability across various LLM backbones and tasks highlight its potential as a versatile alignment technique. We will release the code to foster research in this area. Jingyuan Yang 0008, Wei Peng 0011 |
ICLR | 3 |
| 2025 | Latent Feature Activation Steering for Enhancing Semantic Consistency in Large Language Models
Jingyuan Yang 0008, Rongjun Li, Zhiyong Feng 0002, Wei Peng 0011 |
NLPCC (2) | 6 |
| 2025 | Gradient Co-occurrence Analysis for Detecting Unsafe Prompts in Large Language Models
Jingyuan Yang 0008, Rongjun Li, Zhiyong Feng 0002, Wei Peng 0011 |
NLPCC (2) | 7 |
| 2024 | Learning from Failure: Improving Meeting Summarization without Good SamplesabstractExisting methods aligning language models with various human needs are reliant heavily on high-quality and task-specific data. However, industrial deployment of task-specific language models often encounter challenges in the availability of appropriate training samples. Taking meeting summarization for instance, public datasets are scarce, and private corpora are also hard to obtain due to privacy issues or resource-demanding annotation. To improve meeting summarization in the absence of positively-rated (i.e., ``good'') samples, we propose Score Tuning, a cold start tuning framework that leverages bad samples of distinguishable degrees to incrementally enhance the performance of summary generation without an initial presence of good samples. Our method utilizes asynchronous and numerical human feedback that measure the quality of generated summaries. Formulating data into triplets of (transcript, summary, score), our approach instructs a pre-trained model to learn the association between summary qualities and human-rated scores and hence to generate better summaries corresponding to higher scores. The experiment results show that our method is effective in improving meeting summarization on both English and Chinese corpora while requiring less annotated data and training resources compared to existing alignment methods. Additionally, we also preliminarily explore the transferability of our approach in machine translation tasks and demonstrate its potential for future development and usage in other domains. Xiutian Zhao, Wei Peng 0011 |
AAAI | 3 |
| 2024 | Self-Distillation Regularized Connectionist Temporal Classification Loss for Text Recognition: A Simple Yet Effective ApproachabstractText recognition methods are gaining rapid development. Some advanced techniques, e.g., powerful modules, language models, and un- and semi-supervised learning schemes, consecutively push the performance on public benchmarks forward. However, the problem of how to better optimize a text recognition model from the perspective of loss functions is largely overlooked. CTC-based methods, widely used in practice due to their good balance between performance and inference speed, still grapple with accuracy degradation. This is because CTC loss emphasizes the optimization of the entire sequence target while neglecting to learn individual characters. We propose a self-distillation scheme for CTC-based model to address this issue. It incorporates a framewise regularization term in CTC loss to emphasize individual supervision, and leverages the maximizing-a-posteriori of latent alignment to solve the inconsistency problem that arises in distillation between CTC-based models. We refer to the regularized CTC loss as Distillation Connectionist Temporal Classification (DCTC) loss. DCTC loss is module-free, requiring no extra parameters, longer inference lag, or additional training data or phases. Extensive experiments on public benchmarks demonstrate that DCTC can boost text recognition model accuracy by up to 2.6%, without any of these drawbacks. Ziyin Zhang, Ning Lu 0003, Minghui Liao, Yongshuai Huang, Cheng Li 0040, Wei Peng 0011 |
AAAI | 7 |
| 2024 | Contextual Modeling for Document-level ASR Error CorrectionabstractContextual information, including the sentences in the same document and in other documents of the dataset, plays a crucial role in improving the accuracy of document-level ASR Error Correction (AEC), while most previous works ignore this. In this paper, we propose a context-aware method that utilizes a k-Nearest Neighbors (kNN) approach to enhance the AEC model by retrieving a datastore containing contextual information. We conduct experiments on two English and two Chinese datasets, and the results demonstrate that our proposed model can effectively utilize contextual information to improve document-level AEC. Furthermore, the context information from the whole dataset provides even better results. Xunjian Yin, Xiaojun Wan 0001, Wei Peng 0011, Rongjun Li, Jingyuan Yang 0008, Yanquan Zhou |
LREC/COLING | 4 |
| 2024 | Align Before Adapt: Leveraging Entity-to-Region Alignments for Generalizable Video Action RecognitionabstractLarge-scale visual-language pre-trained models have achieved significant success in various video tasks. However, most existing methods follow an “adapt then align” paradigm, which adapts pre-trained image encoders to model video-level representations and utilizes one-hot or text embedding of the action labels for supervision. This paradigm overlooks the challenge of mapping from static images to complicated activity concepts. In this paper, we propose a novel “Align before Adapt” (ALT) paradigm. Prior to adapting to video representation learning, we exploit the entity-to-region alignments for each frame. The alignments are fulfilled by matching the region-aware image embeddings to an offline-constructed text corpus. With the aligned entities, we feed their text embeddings to a transformer-based video adapter as the queries, which can help extract the semantics of the most important entities from a video to a vector. This paradigm reuses the visual-language alignment of VLP during adaptation and tries to explain an action by the underlying entities. This helps understand actions by bridging the gap with complex activity semantics, particularly when facing unfamiliar or unseen categories. ALT demonstrates competitive performance while maintaining remarkably low computational costs. In fully supervised experiments, it achieves 88.1 % top-1 accuracy on Kinetics-400 with only 4947 GFLOPs. Moreover, ALT outperforms the previous state-of-the-art methods in both zero-shot and fewshot experiments, emphasizing its superior generalizability across various learning scenarios. Yifei Chen 0010, Dapeng Chen, Ruijin Liu, Wenyuan Xue, Wei Peng 0011 |
CVPR | 6 |
| 2024 | Watch Every Step! LLM Agent Learning via Iterative Step-level Process RefinementabstractLarge language model agents have exhibited exceptional performance across a range of complex interactive tasks.Recent approaches have utilized tuning with expert trajectories to enhance agent performance, yet they primarily concentrate on outcome rewards, which may lead to errors or suboptimal actions due to the absence of process supervision signals.In this paper, we introduce the Iterative step-level Process Refinement (IPR) framework, which provides detailed step-by-step guidance to enhance agent training.Specifically, we adopt the Monte Carlo method to estimate step-level rewards.During each iteration, the agent explores along the expert trajectory and generates new actions.These actions are then evaluated against the corresponding step of expert trajectory using step-level rewards.Such comparison helps identify discrepancies, yielding contrastive action pairs that serve as training data for the agent.Our experiments on three complex agent tasks demonstrate that our framework outperforms a variety of strong baselines.Moreover, our analytical findings highlight the effectiveness of IPR in augmenting action efficiency and its applicability to diverse models † . Weimin Xiong, Yifan Song 0002, Xiutian Zhao, Cheng Li 0040, Wei Peng 0011, Sujian Li |
EMNLP | 8 |
| 2024 | An Electoral Approach to Diversify LLM-based Multi-Agent Collective Decision-MakingabstractModern large language models (LLMs) have exhibited cooperative synergy on complex tasksolving, and collective decision-making (CDM) is a pivotal component in LLM-based multiagent collaboration frameworks.Our survey on 52 recent such systems uncovers a severe lack of diversity, with a heavy reliance on dictatorial and plurality voting for CDM.Through the lens of social choice theory, we scrutinize widelyadopted CDM methods and identify their limitations.To enrich current landscape of LLMbased CDM, we present GEDI, an electoral CDM module that incorporates various ordinal preferential voting mechanisms.Our empirical case study across three benchmarks shows that the integration of certain CDM methods can markedly improve the reasoning capabilities and robustness of some leading LLMs, all without requiring intricate system designs.Additionally, we find that some CDM mechanisms generate positive synergies even with as few as three agents.The voting-based methods also demonstrate robustness against single points of failure, as well as diversity in terms of hitrate@k and subject-wise impacts. 1 Xiutian Zhao, Wei Peng 0011 |
EMNLP | 3 |
| 2024 | Cross Modal Training for ASR Error Correction with Contrastive LearningabstractASR Error Correction (AEC) aims to post-process the output of ASR systems and further reduce the word error rate. In this paper, we propose a cross-modal training framework with contrastive learning on the AEC task. This framework enables a shared encoder-decoder model to learn text, pinyin (phoneme1) and audio information simultaneously, which is trained by three subtasks: text correction, pinyin to text and ASR. On this basis, we introduce contrastive learning loss to shrink the distance between the three modalities and construct a unified representation. Experiments2on four AEC datasets show that our method effectively corrects a large number of ASR errors to state-of-the-art levels. Xiaojun Wan 0001, Wei Peng 0011, Rongjun Li, Jingyuan Yang 0008, Yanquan Zhou |
ICASSP | 3 |
| 2024 | Assessing Factual Reliability of Large Language Model KnowledgeabstractWeixuan Wang, Barry Haddow, Alexandra Birch, Wei Peng. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Barry Haddow, Alexandra Birch, Wei Peng 0011 |
NAACL-HLT | 4 |
| 2024 | A Neuroinspired Contrast Mechanism enables Few-Shot Object Detection
Lingxiao Yang, Dapeng Chen, Yifei Chen 0010, Wei Peng 0011, Xiaohua Xie |
Pattern Recognit. | 4 |
| 2024 | Reverse Backdoor Distillation: Towards Online Backdoor Attack Detection for Deep Neural Network ModelsabstractThe backdoor attack on deep neural network models implants malicious data patterns in a model to induce attacker-desirable behaviors. Existing defense methods fall into the online and offline categories, in which the offline models achieve state-of-the-art detection rates but are restricted by heavy computation overhead. In contrast, their more deployable online counterparts lack the means to detect source-specific backdoors with large sizes. This work proposes a new online backdoor detection method—Reverse Backdoor Distillation (RBD) to handle issues associated with source-specific and source-agnostic backdoor attacks. RBD, designed with the novel perspective of distilling instead of erasing backdoor knowledge, is a complementary backdoor detection methodology that can be used in conjunction with other online backdoor defenses. Considering the fact that trigger data will cause overwhelming neuron activation while clean data will not, RBD distills backdoor attack pattern knowledge from a suspicious model to create a shadow model, which is subsequently deployed online along with the original model in scope to predict a backdoor attack. We extensively evaluate RBD on several datasets (MNIST, GTSRB, CIFAR-10) with diverse model architectures and trigger patterns. RBD outperforms online benchmarks in all experimental settings. Notably, RBD demonstrates superior capability in detecting source-specific attacks, where comparison methods fail, underscoring the effectiveness of our proposed technique. Moreover, RBD achieves a computational savings of at least 97%. Zeming Yao, Hangtao Zhang, Yicheng Guo, Xin Tian 0015, Wei Peng 0011, Yi Zou 0001, Leo Yu Zhang, Chao Chen 0015 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2023 | Improving Table Structure Recognition with Visual-Alignment Sequential Coordinate ModelingabstractTable structure recognition aims to extract the logical and physical structure of unstructured table images into a machine-readable format. The latest end-to-end image-to-text approaches simultaneously predict the two structures by two decoders, where the prediction of the physical structure (the bounding boxes of the cells) is based on the representation of the logical structure. However, the previous methods struggle with imprecise bounding boxes as the logical representation lacks local visual information. To address this issue, we propose an end-to-end sequential modeling framework for table structure recognition called VAST. It contains a novel coordinate sequence decoder triggered by the representation of the non-empty cell from the logical structure decoder. In the coordinate sequence decoder, we model the bounding box coordinates as a language sequence, where the left, top, right and bottom coordinates are decoded sequentially to leverage the inter-coordinate dependency. Furthermore, we propose an auxiliary visual-alignment loss to enforce the logical representation of the non-empty cells to contain more local visual details, which helps produce better cell bounding boxes. Extensive experiments demonstrate that our proposed method can achieve state-of-the-art results in both logical and physical structure recognition. The ablation study also validates that the proposed coordinate sequence decoder and the visual-alignment loss are the keys to the success of our method. Yongshuai Huang, Ning Lu 0003, Dapeng Chen, Zecheng Xie, Shenggao Zhu, Liangcai Gao, Wei Peng 0011 |
CVPR | 8 |
| 2023 | PROSE: A Pronoun Omission Solution for Chinese-English Spoken Language TranslationabstractNeural Machine Translation (NMT) systems encounter a significant challenge when translating a pro-drop ('pronoun-dropping') language (e.g., Chinese) to a non-pro-drop one (e.g., English), since the pro-drop phenomenon demands NMT systems to recover omitted pronouns.This unique and crucial task, however, lacks sufficient datasets for benchmarking.To bridge this gap, we introduce PROSE, a new benchmark featured in diverse pro-drop instances for document-level Chinese-English spoken language translation.Furthermore, we conduct an in-depth investigation of the prodrop phenomenon in spoken Chinese on this dataset, reconfirming that pro-drop reduces the performance of NMT systems in Chinese-English translation.To alleviate the negative impact introduced by pro-drop, we propose Mention-Aware Semantic Augmentation, a novel approach that leverages the semantic embedding of dropped pronouns to augment training pairs.Results from the experiments on four Chinese-English translation corpora show that our proposed method outperforms existing methods regarding omitted pronoun retrieval and overall translation quality. Xiutian Zhao, Yanghui Li, Wei Peng 0011 |
EMNLP | 4 |
| 2023 | M³Seg: A Maximum-Minimum Mutual Information Paradigm for Unsupervised Topic Segmentation in ASR TranscriptsabstractTopic segmentation aims to detect topic boundaries and split automatic speech recognition transcriptions (e.g., meeting transcripts) into segments that are bounded by thematic meanings.In this work, we propose M 3 Seg, a novel Maximum-Minimum Mutual information paradigm for linear topic segmentation without using any parallel data.Specifically, by employing sentence representations provided by pre-trained language models, M 3 Seg first learns a region-based segment encoder based on the maximization of mutual information between the global segment representation and the local contextual sentence representation.Secondly, an edge-based boundary detection module aims to segment the whole by topics based on minimizing the mutual information between different segments.Experiment results on two public datasets demonstrate the effectiveness of M 3 Seg, which outperform the state-of-the-art methods by a significant (18%-37% improvement) margin. Xiutian Zhao, Yanghui Li, Wei Peng 0011 |
EMNLP | 4 |
| 2023 | ORCHID: A Chinese Debate Corpus for Target-Independent Stance Detection and Argumentative Dialogue SummarizationabstractDialogue agents have been receiving increasing attention for years, and this trend has been further boosted by the recent progress of large language models (LLMs).Stance detection and dialogue summarization are two core tasks of dialogue agents in application scenarios that involve argumentative dialogues.However, research on these tasks is limited by the insufficiency of public datasets, especially for non-English languages.To address this language resource gap in Chinese, we present OR-CHID (Oral Chinese Debate), the first Chinese dataset for benchmarking target-independent stance detection and debate summarization.Our dataset consists of 1,218 real-world debates that were conducted in Chinese on 476 unique topics, containing 2,436 stance-specific summaries and 14,133 fully annotated utterances.Besides providing a versatile testbed for future research, we also conduct an empirical study on the dataset and propose an integrated task.The results show the challenging nature of the dataset and suggest a potential of incorporating stance detection in summarization for argumentative dialogue.1 Xiutian Zhao, Wei Peng 0011 |
EMNLP | 3 |
| 2023 | Knowledge-Augmented Frame Semantic Parsing with Hybrid Prompt-TuningabstractFrame semantics-based approaches have been widely used in semantic parsing tasks and have become mainstream. It remains challenging to disambiguate frame representations evoked by target lexical units under different contexts. Pre-trained Language Models (PLMs) have been used in semantic parsing and significantly improve the accuracy of neural parsers. However, the PLMs-based approaches tend to favor collocated patterns presented in the training data, leading to inaccurate outcomes. The intuition here is to design a mechanism to optimally use knowledge captured in semantic frames in conjunction with PLMs to disambiguate frames. We propose a novel Knowledge-Augmented Frame Semantic Parsing Architecture (KAF-SPA) to enhance semantic representation by incorporating accurate frame knowledge into PLMs during frame semantic parsing. Specifically, a Memory-based Knowledge Extraction Module (MKEM) is devised to select accurate frame knowledge and construct the continuous templates in the high dimensional vector space. Moreover, we design a Task-oriented Knowledge Probing Module (TKPM) using hybrid prompts (in terms of continuous and discrete prompts) to incorporate the selected knowledge into the PLMs and adapt PLMs to the tasks of frame and argument identification. Experimental results on two public FrameNet datasets demonstrate that our method significantly outperforms strong baselines (by more than +3% in F1), achieving state-of-art results on the current benchmark. Ablation studies verify the effectiveness of KAF-SPA. Yajing Sun, Jingyuan Yang 0008, Wei Peng 0011 |
ICASSP | 4 |
| 2023 | Video Action Recognition with Attentive Semantic UnitsabstractVisual-Language Models (VLMs) have significantly advanced video action recognition. Supervised by the semantics of action labels, recent works adapt the visual branch of VLMs to learn video representations. Despite the effectiveness proved by these works, we believe that the potential of VLMs has yet to be fully harnessed. In light of this, we exploit the semantic units (SU) hiding behind the action labels and leverage their correlations with fine-grained items in frames for more accurate action recognition. SUs are entities extracted from the language descriptions of the entire action set, including body parts, objects, scenes, and motions. To further enhance the alignments between visual contents and the SUs, we introduce a multi-region attention module (MRA) to the visual branch of the VLM. The MRA allows the perception of region-aware visual features beyond the original global feature. Our method adaptively attends to and selects relevant SUs with visual features of frames. With a cross-modal decoder, the selected SUs serve to decode spatiotemporal video representations. In summary, the SUs as the medium can boost discriminative ability and transferability. Specifically, in fully-supervised learning, our method achieved 87.8% top-1 accuracy on Kinetics-400. In K=2 few-shot experiments, our method surpassed the previous state-of-the-art by +7.1% and +15.0% on HMDB-51 and UCF-101, respectively. Yifei Chen 0010, Dapeng Chen, Ruijin Liu, Wei Peng 0011 |
ICCV | 5 |
| 2023 | Video Summarization Leveraging Multimodal Information for Presentations
Dapeng Chen, Rongjun Li, Wenyuan Xue, Wei Peng 0011 |
INTERSPEECH | 5 |
| 2023 | PBFormer: Capturing Complex Scene Text Shape with Polynomial Band TransformerabstractWe present PBFormer, an efficient yet powerful scene text detector that unifies the transformer with a novel text shape representation Polynomial Band (PB). The representation has four polynomial curves to fit a text's top, bottom, left, and right sides, which can capture a text with a complex shape by varying polynomial coefficients. PB has appealing features compared with conventional representations: 1) It can model different curvatures with a fixed number of parameters, while polygon-points-based methods need to utilize a different number of points. 2) It can distinguish adjacent or overlapping texts as they have apparent different curve coefficients, while segmentation-based or points-based methods suffer from adhesive spatial positions. PBFormer combines the PB with the transformer, which can directly generate smooth text contours sampled from predicted curves without interpolation. A parameter-free cross-scale pixel attention (CPA) module is employed to highlight the feature map of a suitable scale while suppressing the other feature maps. The simple operation can help detect small-scale texts and is compatible with the one-stage DETR framework, where no postprocessing exists for NMS. Furthermore, PBFormer is trained with a shape-contained loss, which not only enforces the piecewise alignment between the ground truth and the predicted curves but also makes curves' position and shapes consistent with each other. Without bells and whistles about text pre-training, our method is superior to the previous state-of-the-art text detectors on the arbitrary-shaped text datasets. Codes will be public. Ruijin Liu, Ning Lu 0003, Dapeng Chen, Cheng Li 0040, Zejian Yuan, Wei Peng 0011 |
ACM Multimedia | 6 |
| 2023 | An empirical study of cyclical learning rate on neural machine translationabstractAbstract In training deep learning networks, the optimizer and related learning rate are often used without much thought or with minimal tuning, even though it is crucial in ensuring a fast convergence to a good quality minimum of the loss function that can also generalize well on the test dataset. Drawing inspiration from the successful application of cyclical learning rate policy to computer vision tasks, we explore how cyclical learning rate can be applied to train transformer-based neural networks for neural machine translation. From our carefully designed experiments, we show that the choice of optimizers and the associated cyclical learning rate policy can have a significant impact on the performance. In addition, we establish guidelines when applying cyclical learning rates to neural machine translation tasks. Choon Meng Lee, Jianfeng Liu 0004, Talha Çolakoglu, Wei Peng 0011 |
Nat. Lang. Eng. | 5 |
| 2022 | Cross-lingual Feature Extraction from Monolingual Corpora for Low-resource Unsupervised Bilingual Lexicon InductionabstractDespite their progress in high-resource language settings, unsupervised bilingual lexicon induction (UBLI) models often fail on corpora with low-resource distant language pairs due to insufficient initialization. In this work, we propose a cross-lingual feature extraction (CFE) method to learn the cross-lingual features from monolingual corpora for low-resource UBLI, enabling representations of words with the same meaning leveraged by the initialization step. By integrating cross-lingual representations with pre-trained word embeddings in a fully unsupervised initialization on UBLI, the proposed method outperforms existing state-of-the-art methods on low-resource language pairs (EN-VI, EN-TH, EN-ZH, EN-JA). The ablation study also proves that the learned cross-lingual features can enhance the representational ability and robustness of the existing embedding model. Hailong Cao, Tiejun Zhao, Wei Peng 0011 |
COLING | 5 |
| 2022 | ASR Error Correction with Constrained Decoding on Operation Prediction
Jingyuan Yang 0008, Rongjun Li, Wei Peng 0011 |
INTERSPEECH | 3 |
| 2022 | Cross-Modal Retrieval with Heterogeneous Graph EmbeddingabstractConventional methods address the cross-modal retrieval problem by projecting the multi-modal data into a shared representation space. Such a strategy will inevitably lose the modality-specific information, leading to decreased retrieval accuracy. In this paper, we propose heterogeneous graph embeddings to preserve more abundant cross-modal information. The embedding from one modality will be compensated with the aggregated embeddings from the other modality. In particular, a self-denoising tree search is designed to reduce the "label noise" problem, making the heterogeneous neighborhood more semantically relevant. The dual-path aggregation tackles the "modality imbalance" problem, giving each sample comprehensive dual-modality information. The final heterogeneous graph embedding is obtained by feeding the aggregated dual-modality features to the cross-modal self-attention module. Experiments conducted on cross-modality person re-identification and image-text retrieval task validate the superiority and generality of the proposed method. Dapeng Chen, Lin Wu 0001, Harry Qin, Wei Peng 0011 |
ACM Multimedia | 6 |
| 2022 | FNeVR: Neural Volume Rendering for Face AnimationabstractFace animation, one of the hottest topics in computer vision, has achieved a promising performance with the help of generative models. However, it remains a critical challenge to generate identity preserving and photo-realistic images due to the sophisticated motion deformation and complex facial detail modeling. To address these problems, we propose a Face Neural Volume Rendering (FNeVR) network to fully explore the potential of 2D motion warping and 3D volume rendering in a unified framework. In FNeVR, we design a 3D Face Volume Rendering (FVR) module to enhance the facial details for image rendering. Specifically, we first extract 3D information with a well designed architecture, and then introduce an orthogonal adaptive ray-sampling module for efficient rendering. We also design a lightweight pose editor, enabling FNeVR to edit the facial pose in a simple yet effective way. Extensive experiments show that our FNeVR obtains the best overall quality and performance on widely used talking-head benchmarks. Bohan Zeng, Hong Li 0016, Xuhui Liu, Jianzhuang Liu, Dapeng Chen, Wei Peng 0011, Baochang Zhang 0001 |
NeurIPS | 7 |
| 2021 | Robustness Testing of Language Understanding in Task-Oriented DialogabstractJiexi Liu, Ryuichi Takanobu, Jiaxin Wen, Dazhen Wan, Hongguang Li, Weiran Nie, Cheng Li, Wei Peng, Minlie Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jiexi Liu 0002, Ryuichi Takanobu, Jiaxin Wen, Dazhen Wan, Weiran Nie, Cheng Li 0040, Wei Peng 0011, Minlie Huang |
ACL/IJCNLP (1) | 8 |
| 2021 | Neural Machine Translation with Heterogeneous Topic Knowledge EmbeddingsabstractNeural Machine Translation (NMT) has shown a strong ability to utilize local context to disambiguate the meaning of words.However, it remains a challenge for NMT to leverage broader context information like topics.In this paper, we propose heterogeneous ways of embedding topic information at the sentence level into an NMT model to improve translation performance.Specifically, the topic information can be incorporated as pre-encoder topic embedding, post-encoder topic embedding, and decoder topic embedding to increase the likelihood of selecting target words from the same topic of the source sentence.Experimental results show that NMT models with the proposed topic knowledge embedding outperform the baselines on the English → German and English → French translation tasks. Wei Peng 0011, Meng Zhang 0019, Qun Liu 0001 |
EMNLP (1) | 2 |
| 2021 | Coreference Augmentation for Multi-Domain Task-Oriented Dialogue State TrackingabstractDialogue State Tracking (DST), which is the process of inferring user goals by estimating belief states given the dialogue history, plays a critical role in task-oriented dialogue systems.A coreference phenomenon observed in multi-turn conversations is not addressed by existing DST models, leading to suboptimal performances.In this paper, we propose Coreference Dialogue State Tracker (CDST) that explicitly models the coreference feature.In particular, at each turn, the proposed model jointly predicts the coreferred domain-slot pair and extracts the coreference values from the dialogue context.Experimental results on MultiWOZ 2.1 dataset show that the proposed model achieves the state-of-the-art joint goal accuracy of 56.47%. Chongxuan Huang, Wei Peng 0011 |
Interspeech | 3 |
| 2021 | MultiWOZ 2.3: A Multi-domain Task-Oriented Dialogue Dataset Enhanced with Annotation Corrections and Co-Reference Annotation
Ryuichi Takanobu, Yixin Lian, Chongxuan Huang, Dazhen Wan, Wei Peng 0011, Minlie Huang |
NLPCC (2) | 7 |
| 2019 | Survey on SDN based network intrusion detection system using machine learning approaches
Nasrin Sultana, Naveen K. Chilamkurti, Wei Peng 0011, Rabei Alhadad |
Peer-to-Peer Netw. Appl. | 3 |
| 2017 | Roles of policy settings in distributed generation with battery storageabstractDistributed Generation (DG) is a sustainable alternative energy paradigm that allows flexible customer-participated demand response management, however when coupled with battery storage in a carbon costed policy setting true reduction of greenhouse gas emissions may not necessarily be rewarded. This paper examines the role of policy settings using an established multi-agent simulation framework that captures emerging complex responses that originate from individual household energy use behaviors. Case studies demonstrate with uninformed policy settings being chosen, undesirable over-generation may cause technical issues with unwanted energy profile responses as well as undesirable over-investment in the wrong electricity assets may cause increased electricity costs for households. Peter Sokolowski, Wei Peng 0011, Ragini Patel, Xinghuo Yu 0001 |
IECON | 2 |
| 2009 | Putting Simple Hierarchy into Ant Foraging: Cluster-Based Soft-BotsabstractThis paper revisits a traditional Ant Foraging algorithm and proposes a Cluster-based Softbots algorithm to address the performance issues caused by constraints of random autonomous search featured in most swarm intelligence-based algorithms. A simple hierarchy is introduced to regulate the unfolding of dynamically changing swarm-like behaviors. Comparative experiments for Ant Foraging and the proposed Cluster-based Softbots are described. The results demonstrate that Softbots have significant comparative advantages over a traditional Ant Foraging algorithm on the benchmark criteria in the presented experimental settings. It is shown that Softbots are more suitable for resource-lean search circumstances whereas not many individual agents can be allocated. Wei Peng 0011, Qingmai Wang, Xinghuo Yu 0001 |
NSS | 1 |
| 2009 | Understanding behaviors of a constructive memory agent: A markov chain analysis
John S. Gero, Wei Peng 0011 |
Knowl. Based Syst. | 2 |
| 2006 | Using a Constructive Interactive Activation and Competition Neural Network to Construct a Situated Agent's Experience
Wei Peng 0011, John S. Gero |
PRICAI | 1 |