Hongyu Gong

dblp:163/7318 · DBLP profile ↗
← Back
36ranked-venue papers
13as first author
23since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 9 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 10 since 2021Computer networks · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
abstract
Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framework that uses both audio and visual cues to track the target speaker, predict turn-taking, and generate coherent responses. By combining acoustic tokenization with multi-task, multi-stage training on monadic, synthetic, and real audio-visual dialogue datasets, AV-Dialog achieves robust streaming transcription, semantically grounded turn-boundary detection and accurate responses, resulting in a natural conversational flow. Experiments show that AV-Dialog outperforms audio-only models under interference, reducing transcription errors, improving turn-taking prediction, and enhancing human-rated dialogue quality. These results highlight the power of seeing as well as hearing for speaker-aware interaction, paving the way for {spoken} dialogue agents that perform {robustly} in real-world, noisy environments.
Tuochao Chen, Bandhav Veluri, Hongyu Gong, Shyamnath Gollakota
ACL (1)3
2025 AV-Flow: Transforming Text to Audio-Visual Human-Like Interactions
abstract
We introduce AV-Flow, an audio-visual generative model that animates photo-realistic 4D talking avatars given only text input. In contrast to prior work that assumes an existing speech signal, we synthesize speech and vision jointly. We demonstrate human-like speech synthesis, synchronized lip motion, lively facial expressions and head pose; all generated from just text characters. The core premise of our approach lies in the architecture of our two parallel diffusion transformers. Intermediate highway connections ensure communication between the audio and visual modalities, and thus, synchronized speech intonation and facial dynamics (e.g., eyebrow motion). Our model is trained with flow matching, leading to expressive results and fast inference. In case of dyadic conversations, AV-Flow produces an always-on avatar, that actively listens and reacts to the audio-visual input of a user. Through extensive experiments, we show that our method outperforms prior work, synthesizing natural-looking 4D talking avatars. Project page: https://aggelinacha.github.io/AV-Flow/
Aggelina Chatziagapi, Louis-Philippe Morency, Hongyu Gong, Michael Zollhöfer, Dimitris Samaras, Alexander Richard
ICCV3
2025 DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
Heng-Jui Chang, Hongyu Gong, Changhan Wang, James R. Glass, Yu-An Chung
INTERSPEECH2
2024 Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
abstract
Despite broad interest in modeling spoken dialogue agents, most approaches are inherently "half-duplex" -restricted to turn-based interaction with responses requiring explicit prompting by the user or implicit tracking of interruption or silence events.Human dialogue, by contrast, is "full-duplex" allowing for rich synchronicity in the form of quick and dynamic turn-taking, overlapping speech, and backchanneling.Technically, the challenge of achieving full-duplex dialogue with LLMs lies in modeling synchrony as pre-trained LLMs do not have a sense of "time".To bridge this gap, we propose Synchronous LLMs for fullduplex spoken dialogue modeling.We design a novel mechanism to integrate time information into Llama3-8b so that they run synchronously with the real-world clock.We also introduce a training recipe that uses 212k hours of synthetic spoken dialogue data generated from text dialogue data to create a model that generates meaningful and natural spoken dialogue, with just 2k hours of real-world spoken dialogue data.Synchronous LLMs outperform state-of-the-art in dialogue meaningfulness while maintaining naturalness.Finally, we demonstrate the model's ability to participate in full-duplex dialogue by simulating interaction between two agents trained on different datasets, while considering Internet-scale latencies of up to 240ms.
Bandhav Veluri, Benjamin N. Peloquin, Bokai Yu, Hongyu Gong, Shyamnath Gollakota
EMNLP4
2024 Investigating Decoder-only Large Language Models for Speech-to-text Translation
Chao-Wei Huang, Hongyu Gong, Hirofumi Inaguma, Ilia Kulikov, Ruslan Mavlyutov, Sravya Popuri
INTERSPEECH3
2023 SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations
abstract
Paul-Ambroise Duquenne, Hongyu Gong, Ning Dong, Jingfei Du, Ann Lee, Vedanuj Goswami, Changhan Wang, Juan Pino, Benoît Sagot, Holger Schwenk. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Paul-Ambroise Duquenne, Hongyu Gong, Jingfei Du, Ann Lee 0001, Vedanuj Goswami, Changhan Wang, Juan Pino 0001, Benoît Sagot, Holger Schwenk
ACL (1)2
2023 Named Entity Detection and Injection for Direct Speech Translation
abstract
In a sentence, certain words are critical for its semantic. Among them, named entities (NEs) are notoriously challenging for neural models. Despite their importance, their accurate handling has been neglected in speech-to-text (S2T) translation research, and recent work has shown that S2T models perform poorly for locations and notably person names, whose spelling is challenging unless known in advance. In this work, we explore how to leverage dictionaries of NEs known to likely appear in a given context to improve S2T model outputs. Our experiments show that we can reliably detect NEs likely present in an utterance starting from S2T encoder outputs. Indeed, we demonstrate that the current detection quality is sufficient to improve NE accuracy in the translation with a 31% reduction in person name errors.
Marco Gaido, Yun Tang 0002, Ilia Kulikov, Rongqing Huang, Hongyu Gong, Hirofumi Inaguma
ICASSP5
2023 A Holistic Cascade System, Benchmark, and Human Evaluation Protocol for Expressive Speech-to-Speech Translation
abstract
Expressive speech-to-speech translation (S2ST) aims to transfer prosodic attributes of source speech to target speech while maintaining translation accuracy. Existing research in expressive S2ST is limited, typically focusing on a single expressivity aspect at a time. Likewise, this research area lacks standard evaluation protocols and well-curated benchmark datasets. In this work, we propose a holistic cascade system for expressive S2ST, combining multiple prosody transfer techniques previously considered only in isolation. We curate a benchmark expressivity test set in the TV series domain and explored a second dataset in the audiobook domain. Finally, we present a human evaluation protocol to assess multiple expressive dimensions across speech pairs. Experimental results indicate that bi-lingual annotators can assess the quality of expressive preservation in S2ST systems, and the holistic modeling approach outperforms single-aspect systems. Audio samples can be accessed through our demo webpage: https://facebookresearch.github.io/speech_translation/cascade_expressive_s2st.
Wen-Chin Huang, Benjamin N. Peloquin, Justine Kao, Changhan Wang, Hongyu Gong, Elizabeth Salesky, Yossi Adi, Ann Lee 0001, Peng-Jen Chen
ICASSP5
2023 Improving Speech-to-Speech Translation Through Unlabeled Text
abstract
Direct speech-to-speech translation (S2ST) is among the most challenging problems in the translation paradigm due to the significant scarcity of S2ST data. While effort has been made to increase the data size from unlabeled speech by cascading pretrained speech recognition (ASR), machine translation (MT) and text-to-speech (TTS) models; unlabeled text has remained relatively under-utilized to improve S2ST. We propose an effective way to utilize the massive existing unlabeled text from different languages to create a large amount of S2ST data to improve S2ST performance by applying various acoustic effects to the generated synthetic data. Empirically our method outperforms the state of the art in Spanish-English translation by up to 2 BLEU. Significant gains by the proposed method are demonstrated in extremely low-resource settings for both Spanish-English and Russian-English translations.
Xuan-Phi Nguyen, Sravya Popuri, Changhan Wang, Yun Tang 0002, Ilia Kulikov, Hongyu Gong
ICASSP6
2023 Pre-training for Speech Translation: CTC Meets Optimal Transport
abstract
The gap between speech and text modalities is a major challenge in speech-to-text translation (ST). Different methods have been proposed to reduce this gap, but most of them require architectural changes in ST training. In this work, we propose to mitigate this issue at the pre-training stage, requiring no change in the ST model. First, we show that the connectionist temporal classification (CTC) loss can reduce the modality gap by design. We provide a quantitative comparison with the more common cross-entropy loss, showing that pre-training with CTC consistently achieves better final ST accuracy. Nevertheless, CTC is only a partial solution and thus, in our second contribution, we propose a novel pre-training method combining CTC and optimal transport to further reduce this gap. Our method pre-trains a Siamese-like model composed of two encoders, one for acoustic inputs and the other for textual inputs, such that they produce representations that are close to each other in the Wasserstein space. Extensive experiments on the standard CoVoST-2 and MuST-C datasets show that our pre-training method applied to the vanilla encoder-decoder Transformer achieves state-of-the-art performance under the no-external-data setting, and performs on par with recent strong multi-task learning systems trained with external data. Finally, our method can also be applied on top of these multi-task systems, leading to further improvements for these models.
Phuong-Hang Le, Hongyu Gong, Changhan Wang, Juan Pino 0001, Benjamin Lecouteux, Didier Schwab
ICML2
2023 Exploration on HuBERT with Multiple Resolution
Jiatong Shi, Yun Tang 0002, Hirofumi Inaguma, Hongyu Gong, Juan Pino 0001, Shinji Watanabe 0001
INTERSPEECH4
2022 Idiomatic Expression Paraphrasing without Strong Supervision
abstract
Idiomatic expressions (IEs) play an essential role in natural language. In this paper, we study the task of idiomatic sentence paraphrasing (ISP), which aims to paraphrase a sentence with an IE by replacing the IE with its literal paraphrase. The lack of large-scale corpora with idiomatic-literal parallel sentences is a primary challenge for this task, for which we consider two separate solutions. First, we propose an unsupervised approach to ISP, which leverages an IE's contextual information and definition and does not require a parallel sentence training set. Second, we propose a weakly supervised approach using back-translation to jointly perform paraphrasing and generation of sentences with IEs to enlarge the small-scale parallel sentence training dataset. Other significant derivatives of the study include a model that replaces a literal phrase in a sentence with an IE to generate an idiomatic expression and a large scale parallel dataset with idiomatic/literal sentence pairs. The effectiveness of the proposed solutions compared to competitive baselines is seen in the relative gains of over 5.16 points in BLEU, over 8.75 points in METEOR, and over 19.57 points in SARI when the generated sentences are empirically validated on a parallel dataset using automatic and manual evaluations. We demonstrate the practical utility of ISP as a preprocessing step in En-De machine translation.
Jianing Zhou, Ziheng Zeng, Hongyu Gong, Suma Bhat
AAAI3
2022 Unified Speech-Text Pre-training for Speech Translation and Recognition
abstract
Yun Tang, Hongyu Gong, Ning Dong, Changhan Wang, Wei-Ning Hsu, Jiatao Gu, Alexei Baevski, Xian Li, Abdelrahman Mohamed, Michael Auli, Juan Pino. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yun Tang 0002, Hongyu Gong, Changhan Wang, Wei-Ning Hsu, Jiatao Gu, Alexei Baevski, Xian Li 0003, Abdel-rahman Mohamed, Michael Auli, Juan Pino 0001
ACL (1)2
2022 T-Modules: Translation Modules for Zero-Shot Cross-Modal Machine Translation
abstract
We present a new approach to perform zeroshot cross-modal transfer between speech and text for translation tasks.Multilingual speech and text are encoded in a joint fixed-size representation space.Then, we compare different approaches to decode these multimodal and multilingual fixed-size representations, enabling zero-shot translation between languages and modalities.All our models are trained without the need of cross-modal labeled translation data.Despite a fixed-size representation, we achieve very competitive results on several text and speech translation tasks.In particular, we outperform the state of the art for zero-shot speech translation on Must-C.We also introduce the first results for zero-shot direct speechto-speech and text-to-speech translation.
Paul-Ambroise Duquenne, Hongyu Gong, Benoît Sagot, Holger Schwenk
EMNLP2
2022 Contrastive Clustering to Mine Pseudo Parallel Data for Unsupervised Translation
Xuan-Phi Nguyen, Hongyu Gong, Yun Tang 0002, Changhan Wang, Philipp Koehn, Shafiq R. Joty
ICLR2
2022 From Start to Finish: Latency Reduction Strategies for Incremental Speech Synthesis in Simultaneous Speech-to-Speech Translation
abstract
Speech-to-speech translation (S2ST) converts input speech to speech in another language. A challenge of delivering S2ST in real time is the accumulated delay between the translation and speech synthesis modules. While recently incremental text-to-speech (iTTS) models have shown large quality improvements, they typically require additional future text inputs to reach optimal performance. In this work, we minimize the initial waiting time of iTTS by adapting the upstream speech translator to generate high-quality pseudo lookahead for the speech synthesizer. After mitigating the initial delay, we demonstrate that the duration of synthesized speech also plays a crucial role on latency. We formalize this as a latency metric and then present a simple yet effective duration-scaling approach for latency reduction. Our approaches consistently reduce latency by 0.2-0.5 second without sacrificing speech translation quality.
Changhan Wang, Hongyu Gong, Xutai Ma, Yun Tang 0002, Juan Pino 0001
INTERSPEECH3
2022 Textless Speech-to-Speech Translation on Real Data
abstract
Ann Lee, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk, Peng-Jen Chen, Changhan Wang, Sravya Popuri, Yossi Adi, Juan Pino, Jiatao Gu, Wei-Ning Hsu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Ann Lee 0001, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk, Peng-Jen Chen, Changhan Wang, Sravya Popuri, Yossi Adi, Juan Pino 0001, Jiatao Gu, Wei-Ning Hsu
NAACL-HLT2
2021 Abusive Language Detection in Heterogeneous Contexts: Dataset Collection and the Role of Supervised Attention
abstract
Abusive language is a massive problem in online social platforms. Existing abusive language detection techniques are particularly ill-suited to comments containing heterogeneous abusive language patterns, i.e., both abusive and non-abusive parts. This is due in part to the lack of datasets that explicitly annotate heterogeneity in abusive language. We tackle this challenge by providing an annotated dataset of abusive language in over 11,000 comments from YouTube. We account for heterogeneity in this dataset by separately annotating both the comment as a whole and the individual sentences that comprise each comment. We then propose an algorithm that uses a supervised attention mechanism to detect and categorize abusive content using multi-task learning. We empirically demonstrate the challenges of using traditional techniques on heterogeneous content and the comparative gains in performance of the proposed approach over state-of-the-art methods.
Hongyu Gong, Alberto Valido, Katherine M. Ingram, Giulia Fanti, Suma Bhat, Dorothy Espelage
AAAI1
2021 WikiMatrix: Mining 135M Parallel Sentences in 1620 Language Pairs from Wikipedia
abstract
Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, Francisco Guzmán. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Holger Schwenk, Vishrav Chaudhary, Hongyu Gong, Francisco Guzmán
EACL4
2021 Multimodal and Multilingual Embeddings for Large-Scale Speech Mining
abstract
We present an approach to encode a speech signal into a fixed-size representation which minimizes the cosine loss with the existing massively multilingual LASER text embedding space. Sentences are close in this embedding space, independently of their language and modality, either text or audio. Using a similarity metric in that multimodal embedding space, we perform mining of audio in German, French, Spanish and English from Librivox against billions of sentences from Common Crawl. This yielded more than twenty thousand hours of aligned speech translations. To evaluate the automatically mined speech/text corpora, we train neural speech translation systems for several languages pairs. Adding the mined data, achieves significant improvements in the BLEU score on the CoVoST2 and the MUST-C test sets with respect to a very competitive baseline. Our approach can also be used to directly perform speech-to-speech mining, without the need to first transcribe or translate the data. We obtain more than one thousand three hundred hours of aligned speech in French, German, Spanish and English. This speech corpus has the potential to boost research in speech-to-speech translation which suffers from scarcity of natural end-to-end training data. All the mined multimodal corpora will be made freely available.
Paul-Ambroise Duquenne, Hongyu Gong, Holger Schwenk
NeurIPS2
2021 Pay Better Attention to Attention: Head Selection in Multilingual and Multi-Domain Sequence Modeling
abstract
Multi-head attention has each of the attention heads collect salient information from different parts of an input sequence, making it a powerful mechanism for sequence modeling. Multilingual and multi-domain learning are common scenarios for sequence modeling, where the key challenge is to maximize positive transfer and mitigate negative interference across languages and domains. In this paper, we find that non-selective attention sharing is sub-optimal for achieving good generalization across all languages and domains. We further propose attention sharing strategies to facilitate parameter sharing and specialization in multilingual and multi-domain sequence modeling. Our approach automatically learns shared and specialized attention heads for different languages and domains. Evaluated in various tasks including speech recognition, text-to-text and speech-to-text translation, the proposed attention sharing strategies consistently bring gains to sequence models built upon multi-head attention. For speech-to-text translation, our approach yields an average of $+2.0$ BLEU over $13$ language directions in multilingual setting and $+2.0$ BLEU over $3$ domains in multi-domain setting.
Hongyu Gong, Yun Tang 0002, Juan Pino 0001, Xian Li 0003
NeurIPS1
2021 Robust Optimization for Multilingual Translation with Imbalanced Data
abstract
Multilingual models are parameter-efficient and especially effective in improving low-resource languages by leveraging crosslingual transfer. Despite recent advance in massive multilingual translation with ever-growing model and data, how to effectively train multilingual models has not been well understood. In this paper, we show that a common situation in multilingual training, data imbalance among languages, poses optimization tension between high resource and low resource languages where the found multilingual solution is often sub-optimal for low resources. We show that common training method which upsamples low resources can not robustly optimize population loss with risks of either underfitting high resource languages or overfitting low resource ones. Drawing on recent findings on the geometry of loss landscape and its effect on generalization, we propose a principled optimization algorithm, Curvature Aware Task Scaling (CATS), which adaptively rescales gradients from different tasks with a meta objective of guiding multilingual training to low-curvature neighborhoods with uniformly low loss for all languages. We ran experiments on common benchmarks (TED, WMT and OPUS-100) with varying degrees of data imbalance. CATS effectively improved multilingual optimization and as a result demonstrated consistent gains on low resources ($+0.8$ to $+2.2$ BLEU) without hurting high resources. In addition, CATS is robust to overparameterization and large batch size training, making it a promising training method for massive multilingual models that truly improve low resource languages.
Xian Li 0003, Hongyu Gong
NeurIPS2
2021 Self-Supervised Euphemism Detection and Identification for Content Moderation
abstract
Fringe groups and organizations have a long history of using euphemisms—ordinary-sounding words with a secret meaning—to conceal what they are discussing. Nowadays, one common use of euphemisms is to evade content moderation policies enforced by social media platforms. Existing tools for enforcing policy automatically rely on keyword searches for words on a "ban list", but these are notoriously imprecise: even when limited to swearwords, they can still cause embarrassing false positives [1]. When a commonly used ordinary word acquires a euphemistic meaning, adding it to a keyword-based ban list is hopeless: consider "pot" (storage container or marijuana?) or "heater" (household appliance or firearm?) The current generation of social media companies instead hire staff to check posts manually, but this is expensive, inhumane, and not much more effective. It is usually apparent to a human moderator that a word is being used euphemistically, but they may not know what the secret meaning is, and therefore whether the message violates policy. Also, when a euphemism is banned, the group that used it need only invent another one, leaving moderators one step behind.This paper will demonstrate unsupervised algorithms that, by analyzing words in their sentence-level context, can both detect words being used euphemistically, and identify the secret meaning of each word. Compared to the existing state of the art, which uses context-free word embeddings, our algorithm for detecting euphemisms achieves 30–400% higher detection accuracies of unlabeled euphemisms in a text corpus. Our algorithm for revealing euphemistic meanings of words is the first of its kind, as far as we are aware. In the arms race between content moderators and policy evaders, our algorithms may help shift the balance in the direction of the moderators.
Wanzheng Zhu, Hongyu Gong, Rohan Bansal, Zachary Weinberg, Nicolas Christin, Giulia Fanti, Suma Bhat
SP2
2020 Recurrent Chunking Mechanisms for Long-Text Machine Reading Comprehension
abstract
In this paper, we study machine reading comprehension (MRC) on long texts, where a model takes as inputs a lengthy document and a question and then extracts a text span from the document as an answer.State-of-the-art models tend to use a pretrained transformer model (e.g., BERT) to encode the joint contextual information of document and question.However, these transformer-based models can only take a fixed-length (e.g., 512) text as its input.To deal with even longer text inputs, previous approaches usually chunk them into equally-spaced segments and predict answers based on each segment independently without considering the information from other segments.As a result, they may form segments that fail to cover the correct answer span or retain insufficient contexts around it, which significantly degrades the performance.Moreover, they are less capable of answering questions that need cross-segment information.We propose to let a model learn to chunk in a more flexible way via reinforcement learning: a model can decide the next segment that it wants to process in either direction.We also employ recurrent mechanisms to enable information to flow across segments.Experiments on three MRC datasets -CoQA, QuAC, and TriviaQA -demonstrate the effectiveness of our proposed recurrent chunking mechanisms: we can obtain segments that are more likely to contain complete answers and at the same time provide sufficient contexts around the ground truth answers for better predictions.
Hongyu Gong, Yelong Shen, Dian Yu 0001, Jianshu Chen, Dong Yu 0001
ACL1
2020 Enriching Word Embeddings with Temporal and Spatial Information
abstract
The meaning of a word is closely linked to sociocultural factors that can change over time and location, resulting in corresponding meaning changes.Taking a global view of words and their meanings in a widely used language, such as English, may require us to capture more refined semantics for use in time-specific or location-aware situations, such as the study of cultural trends or language use.However, popular vector representations for words do not adequately include temporal or spatial information.In this work, we present a model for learning word representation conditioned on time and location.In addition to capturing meaning changes over time and location, we require that the resulting word embeddings retain salient semantic and geometric properties.We train our model on time-and locationstamped corpora, and show using both quantitative and qualitative evaluations that it can capture semantics across time and locations.We note that our model compares favorably with the state-of-the-art for time-specific embedding, and serves as a new benchmark for location-specific embeddings.
Hongyu Gong, Suma Bhat, Pramod Viswanath
CoNLL1
2020 Rich Syntactic and Semantic Information Helps Unsupervised Text Style Transfer
abstract
Text style transfer aims to change an input sentence to an output sentence by changing its text style while preserving the content.Previous efforts on unsupervised text style transfer only use the surface features of words and sentences.As a result, the transferred sentences may either have inaccurate or missing information compared to the inputs.We address this issue by explicitly enriching the inputs via syntactic and semantic structures, from which richer features are then extracted to better capture the original information.Experiments on two text-style-transfer tasks show that our approach improves the content preservation of a strong unsupervised baseline model thereby demonstrating improved transfer performance.
Hongyu Gong, Linfeng Song, Suma Bhat
INLG1
2020 FUSE: Multi-faceted Set Expansion by Coherent Clustering of Skip-Grams
Wanzheng Zhu, Hongyu Gong, Chao Zhang 0014, Jingbo Shang, Suma Bhat, Jiawei Han 0001
ECML/PKDD (3)2
2019 PaRe: A Paper-Reviewer Matching Approach Using a Common Topic Space
abstract
Omer Anjum, Hongyu Gong, Suma Bhat, Wen-Mei Hwu, JinJun Xiong. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Omer Anjum, Hongyu Gong, Suma Bhat, Wen-Mei W. Hwu, Jinjun Xiong
EMNLP/IJCNLP (1)2
2019 Context-Sensitive Malicious Spelling Error Correction
abstract
Misspelled words of the malicious kind work by changing specific keywords and are intended to thwart existing automated applications for cyber-environment control such as harassing content detection on the Internet and email spam detection. In this paper, we focus on malicious spelling correction, which requires an approach that relies on the context and the surface forms of targeted keywords. In the context of two applications-profanity detection and email spam detection-we show that malicious misspellings seriously degrade their performance. We then propose a context-sensitive approach for malicious spelling correction using word embeddings and demonstrate its superior performance compared to state-of-the-art spell checkers.
Hongyu Gong, Suma Bhat, Pramod Viswanath
WWW1
2018 Document Similarity for Texts of Varying Lengths via Hidden Topics
abstract
Measuring similarity between texts is an important task for several applications.Available approaches to measure document similarity are inadequate for document pairs that have non-comparable lengths, such as a long document and its summary.This is because of the lexical, contextual and the abstraction gaps between a long document of rich details and its concise summary of abstract information.In this paper, we present a document matching approach to bridge this gap, by comparing the texts in a common space of hidden topics.We evaluate the matching algorithm on two matching tasks and find that it consistently and widely outperforms strong baselines.We also highlight the benefits of the incorporation of domain knowledge to text matching.
Hongyu Gong, Tarek Sakakini, Suma Bhat, Jinjun Xiong
ACL (1)1
2018 Preposition Sense Disambiguation and Representation
abstract
Prepositions are highly polysemous, and their variegated senses encode significant semantic information.In this paper we match each preposition's left-and right context, and their interplay to the geometry of the word vectors to the left and right of the preposition.Extracting these features from a large corpus and using them with machine learning models makes for an efficient preposition sense disambiguation (PSD) algorithm, which is comparable to and better than state-of-the-art on two benchmark datasets.Our reliance on no linguistic tool allows us to scale the PSD algorithm to a large corpus and learn sensespecific preposition representations.The crucial abstraction of preposition senses as word representations permits their use in downstream applications-phrasal verb paraphrasing and preposition selection-with new state-ofthe-art results.
Hongyu Gong, Jiaqi Mu, Suma Bhat, Pramod Viswanath
EMNLP1
2018 Embedding Syntax and Semantics of Prepositions via Tensor Decomposition
abstract
Hongyu Gong, Suma Bhat, Pramod Viswanath. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Hongyu Gong, Suma Bhat, Pramod Viswanath
NAACL-HLT1
2017 Geometry of Compositionality
abstract
This paper proposes a simple test for compositionality (i.e., literal usage) of a word or phrase in a context-specific way. The test is computationally simple, relying on no external resources and only uses a set of trained word vectors. Experiments show that the proposed method is competitive with state of the art and displays high accuracy in context-specific compositionality detection of a variety of natural language phenomena (idiomaticity, sarcasm, metaphor) for different datasets in multiple languages. The key insight is to connect compositionality to a curious geometric property of word embeddings, which is of independent interest.
Hongyu Gong, Suma Bhat, Pramod Viswanath
AAAI1
2017 Distributed Multicast Tree Construction in Wireless Sensor Networks
abstract
Multicast tree is a key structure for data dissemination from one source to multiple receivers in wireless networks. Minimum length multica modeled as the Steiner tree problem, and is proven to be NP-hard. In this paper, we explore how to efficiently generate minimum length multi wireless sensor networks (WSNs), where only limited knowledge of network topology is available at each node. We design and analyze a simple algorithm, which we call toward source tree (TST), to build multicast trees in WSNs. We show three metrics of TST algorithm, i.e., running and energy efficiency. We prove that its running time is O(√(n log n)), the best among all existing solutions to our best knowledge. We prove that TST tree length is in the same order as Steiner tree, which give a theoretical upper bound and use simulations to show the ratio be only 1.114 when nodes are uniformly distributed. We evaluate energy efficiency in terms of message complexity and the number of forwarding prove that they are both order-optimal. We give an efficient way to construct multicast tree in support of transmission of voluminous data.
Hongyu Gong, Luoyi Fu, Xinzhe Fu, Lutian Zhao, Xinbing Wang
IEEE Trans. Inf. Theory1
2015 A distributed algorithm to construct multicast trees in wireless multi-hop networks
abstract
The minimum-length multicast tree can achieve efficient multicast transmissions and can be formulated as a Steiner Tree. Its construction is non-trivial and has been proven to be NP-hard. In this paper, we combine the design wisdoms in the minimum spanning tree and the shortest path, and propose a new distributed algorithm for constructing an approximate Steiner Tree, particularly applicable to dynamic wireless networks without centralized control. We rigorously prove the performance bounds of our algorithm in terms of tree length, running time and energy consumption. Let m be the multicast group size and n be the network size. We theoretically show that the ratio of our tree length to the minimum value is upper bounded by 2(1 + δ)/√3 (where δ can be any positive value). Simulation results show that this ratio is in fact very close to 1. We also prove that the running time is O( √(mn log m log3n)). The energy consumption is evaluated in terms of message complexity, and is upper bounded by O(n logm). In all, our algorithm achieves the near-optimal tree length, as well as the shortest running time and the lowest message complexity among all solutions we are aware of. We believe our algorithm provides a significant improvement in designing practical routing policies in wireless networks.
Hongyu Gong, Lutian Zhao, Weijie Wu, Xinbing Wang
ICC1
2015 A Distributed Algorithm to Construct Multicast Trees in WSNs: An Approximate Steiner Tree Approach
abstract
Multicast tree is a key structure for data dissemination from one source to multiple receivers in wireless networks. Minimum length multicast tree can be modeled as the Steiner Tree Problem, and is proven to be NP-hard. In this paper, we explore how to efficiently generate minimum length multicast trees in wireless sensor networks (WSNs), where only limited knowledge of network topology is available at each node. We design and analyze a simple and distributed algorithm, which we call Toward Source Tree (TST), to build multicast trees in WSNs. We show three metrics of TST algorithm, i.e., running time, tree length and energy efficiency. We prove that its running time is O(√nlog n), the best among all existing solutions to our best knowledge. We prove that TST tree length is in the same order as Steiner tree, give a theoretical upper bound and use simulations to show the ratio between them is only 1.114 when nodes are uniformly distributed. We evaluate energy efficiency in terms of the number of forwarding nodes in multicast trees, and prove that it is order-optimal. We give an efficient way to construct multicast tree in support of transmission of voluminous data.
Hongyu Gong, Lutian Zhao, Weijie Wu, Xinbing Wang
MobiHoc1