Cícero Nogueira dos Santos

dblp:14/5278 · DBLP profile ↗
← Back
32ranked-venue papers
8as first author
9since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 8 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
20 papers
Language models and text generation · 19% Information extraction and text analysis · 18% Question answering and dialogue systems · 14%
Databases, data mining, and information retrieval
4 papers
Information retrieval · 53% Knowledge graphs · 47%

Topics — the 30 heaviest of 57, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › text summarization
abstractive summarization
0.512021
Improving Factual Consistency of Abstractive Summarization via Question Answering · ACL/IJCNLP (1) 2021
Natural language and speech › Question answering and dialogue systems › robust question answering
ambiguous question answering
0.512021
Answering Ambiguous Questions through Generative Evidence Fusion and Round-Trip Prediction · ACL/IJCNLP (1) 2021
Knowledge, reasoning and agents › Knowledge representation and reasoning › information fusion
evidence combination
0.512021
Answering Ambiguous Questions through Generative Evidence Fusion and Round-Trip Prediction · ACL/IJCNLP (1) 2021
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability › factuality
factual consistency
0.512021
Improving Factual Consistency of Abstractive Summarization via Question Answering · ACL/IJCNLP (1) 2021
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering
0.512021
Generative Context Pair Selection for Multi-hop Question Answering · EMNLP (1) 2021
Machine learning › Graph learning › graph neural network › heterogeneous graph neural network
multi-relational graph neural network
0.512021
Mixed-Curvature Multi-Relational Graph Neural Network for Knowledge Graph Completion · WWW 2021
Machine learning › Representation and self-supervised learning
pre-training
0.512021
Learning Contextual Representations for Semantic Parsing with Generation-Augmented Pre-Training · AAAI 2021
Natural language and speech › Information extraction and text analysis
semantic parsing
0.512021
Learning Contextual Representations for Semantic Parsing with Generation-Augmented Pre-Training · AAAI 2021
Natural language and speech › Language models and text generation › language modeling › language model architecture
sequence-to-sequence model
0.512021
Structured Prediction as Translation between Augmented Natural Languages · ICLR 2021
Machine learning › Probabilistic and Bayesian machine learning
structured prediction
0.512021
Structured Prediction as Translation between Augmented Natural Languages · ICLR 2021
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL
0.512021
Learning Contextual Representations for Semantic Parsing with Generation-Augmented Pre-Training · AAAI 2021
Knowledge graphs
link prediction
0.512021
Mixed-Curvature Multi-Relational Graph Neural Network for Knowledge Graph Completion · WWW 2021
Knowledge graphs › knowledge graph embedding
mixed-curvature embedding
0.512021
Mixed-Curvature Multi-Relational Graph Neural Network for Knowledge Graph Completion · WWW 2021
Natural language and speech › Question answering and dialogue systems › robust question answering
domain adaptation for question answering
0.412020
End-to-End Synthetic Data Generation for Domain Adaptation of Question Answering Systems · EMNLP (1) 2020
Machine learning › Deep learning architectures and training › sequence modeling
generative sequence labeling
0.412020
Augmented Natural Language for Generative Sequence Labeling · EMNLP (1) 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge graph reasoning
knowledge base completion
0.412020
DualTKB: A Dual Learning Bridge between Text and Knowledge Base · EMNLP (1) 2020
Natural language and speech › Language models and text generation › text generation › data-to-text generation
knowledge graph-to-text generation
0.412020
DualTKB: A Dual Learning Bridge between Text and Knowledge Base · EMNLP (1) 2020
Natural language and speech › Information extraction and text analysis
named entity recognition
0.412020
Augmented Natural Language for Generative Sequence Labeling · EMNLP (1) 2020
Natural language and speech › Information extraction and text analysis
sequence labeling
0.412020
Augmented Natural Language for Generative Sequence Labeling · EMNLP (1) 2020
Natural language and speech › Information extraction and text analysis
slot filling
0.412020
Augmented Natural Language for Generative Sequence Labeling · EMNLP (1) 2020
Machine learning › Generative modeling
synthetic data generation
0.412020
End-to-End Synthetic Data Generation for Domain Adaptation of Question Answering Systems · EMNLP (1) 2020
Natural language and speech › Language models and text generation
text generation
0.412020
Learning Implicit Text Generation via Feature Matching · ACL 2020
Information retrieval › question answering
answer selection
0.412020
Beyond [CLS] through Ranking by Generation · EMNLP (1) 2020
Information retrieval › retrieval models
generative retrieval
0.412020
Beyond [CLS] through Ranking by Generation · EMNLP (1) 2020
Knowledge graphs
knowledge graph embedding
0.412020
H2KGAT: Hierarchical Hyperbolic Knowledge Graph Attention Network · EMNLP (1) 2020
Information retrieval
ranking
0.412020
Beyond [CLS] through Ranking by Generation · EMNLP (1) 2020
Information retrieval
retrieval models
0.412020
Beyond [CLS] through Ranking by Generation · EMNLP (1) 2020
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.412019
Wasserstein Barycenter Model Ensembling · ICLR (Poster) 2019
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection
0.412019
Sobolev Independence Criterion · NeurIPS 2019
Machine learning › Generative modeling
generative adversarial network
0.412019
Learning Implicit Generative Models by Matching Perceptual Features · ICCV 2019

Methods — techniques the papers use, named apart from their topics

attention mechanism · 1.1graph neural updater · 1.0trainable curvature · 0.5self-supervised learning · 0.5round-trip prediction · 0.5question answering-based evaluation · 0.5product manifold · 0.5masked language model · 0.5generative model · 0.5generative evidence fusion · 0.5generation model · 0.5augmented natural language · 0.5unlikelihood loss · 0.4hyperbolic embedding · 0.4generative language model · 0.4wasserstein barycenter · 0.4residual learning · 0.3hierarchical recurrent neural network · 0.3
YearPublicationVenuePosition
2024 Memory Augmented Language Models through Mixture of Word Experts
abstract
Cicero Nogueira dos Santos, James Lee-Thorp, Isaac Noble, Chung-Ching Chang, David Uthus. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Cícero Nogueira dos Santos, James Lee-Thorp, Isaac Noble, Chung-Ching Chang, David C. Uthus
NAACL-HLT1
2021 Learning Contextual Representations for Semantic Parsing with Generation-Augmented Pre-Training
abstract
Most recently, there has been significant interest in learning contextual representations for various NLP tasks, by leveraging large scale text corpora to train powerful language models with self-supervised learning objectives, such as Masked Language Model (MLM). Based on a pilot study, we observe three issues of existing general-purpose language models when they are applied in the text-to-SQL semantic parsers: fail to detect the column mentions in the utterances, to infer the column mentions from the cell values, and to compose target SQL queries when they are complex. To mitigate these issues, we present a model pretraining framework, Generation-Augmented Pre-training (GAP), that jointly learns representations of natural language utterance and table schemas, by leveraging generation models to generate high-quality pre-train data. GAP Model is trained on 2 million utterance-schema pairs and 30K utterance-schema-SQL triples, whose utterances are generated by generation models. Based on experimental results, neural semantic parsers that leverage GAP Model as a representation encoder obtain new state-of-the-art results on both Spider and Criteria-to-SQL benchmarks.
Peng Shi 0010, Patrick Ng, Zhiguo Wang 0006, Henghui Zhu, Alexander Hanbo Li, Jun Wang 0122, Cícero Nogueira dos Santos, Bing Xiang
AAAI7
2021 Answering Ambiguous Questions through Generative Evidence Fusion and Round-Trip Prediction
abstract
Yifan Gao, Henghui Zhu, Patrick Ng, Cicero Nogueira dos Santos, Zhiguo Wang, Feng Nan, Dejiao Zhang, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yifan Gao 0001, Henghui Zhu, Patrick Ng, Cícero Nogueira dos Santos, Zhiguo Wang 0006, Feng Nan, Dejiao Zhang, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang
ACL/IJCNLP (1)4
2021 Improving Factual Consistency of Abstractive Summarization via Question Answering
abstract
Feng Nan, Cicero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Kathleen McKeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang, Andrew O. Arnold, Bing Xiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Feng Nan, Cícero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Kathy McKeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang 0006, Andrew O. Arnold, Bing Xiang
ACL/IJCNLP (1)2
2021 Knowledge Graph Representation via Hierarchical Hyperbolic Neural Graph Embedding
abstract
Knowledge graph enhanced information retrieval systems have attracted considerable attention due to their ability to improve performance and provide additional explainability. As the knowledge graphs usually include fruitful facts, they are also good sources of side information. However, recent studies have shown that the usefulness of knowledge graphs depends highly on their representation, e.g., the embeddings of entities and relations. Embedding entities and relations in low-dimensional space is a successful knowledge graph representation solution. Most of the works lie in modeling symmetry/asymmetry/composition/inversion relations but pay less attention to the hierarchical relations. Recent studies have observed the fact that there exist rich semantic hierarchical relations in knowledge graphs such as Freebase (entities are connected in a taxonomic hierarchy) and WordNet (entities are synsets linked together in a hierarchy).To address the above problems, we propose Hierarchical Hyperbolic Neural Graph Embedding (H2E), a new knowledge graph representation approach, which is able to better preserve hierarchical relations. Specifically, the entities/relations representations are learned in a hyperbolic polar embedding space. In a hyperbolic polar embedding space, the entity and relation are modeled as a dual-embedding with modulus embedding part and phase embedding part, enabling the explicitly modeling of two types of hierarchies: inter-level hierarchy and intra-level hierarchy. As the polar embedding is defined i n hyperbolic space, the ability of modeling and inferring hierarchical relations are mutual enhanced. In addition, by noticing the existence of the rich relational context, we propose an attentional neural context aggregation to adaptively integrate the relational context for further enhancing the ability to preserve the hierarchical relations. The empirical study on three benchmark datasets for the link prediction task demonstrates significant performance gains compared to some existing state-of-the-art methods and verifies the effectiveness of the proposed method on hierarchical relations.
Shen Wang 0005, Xiaokai Wei, Cícero Nogueira dos Santos, Zhiguo Wang 0006, Ramesh Nallapati, Andrew O. Arnold, Philip S. Yu
IEEE BigData3
2021 Entity-level Factual Consistency of Abstractive Text Summarization
abstract
Feng Nan, Ramesh Nallapati, Zhiguo Wang, Cicero Nogueira dos Santos, Henghui Zhu, Dejiao Zhang, Kathleen McKeown, Bing Xiang. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Feng Nan, Ramesh Nallapati, Zhiguo Wang 0006, Cícero Nogueira dos Santos, Henghui Zhu, Dejiao Zhang, Kathy McKeown, Bing Xiang
EACL4
2021 Generative Context Pair Selection for Multi-hop Question Answering
abstract
Dheeru Dua, Cicero Nogueira dos Santos, Patrick Ng, Ben Athiwaratkun, Bing Xiang, Matt Gardner, Sameer Singh. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Dheeru Dua, Cícero Nogueira dos Santos, Patrick Ng, Ben Athiwaratkun, Bing Xiang, Matt Gardner 0001, Sameer Singh 0001
EMNLP (1)2
2021 Structured Prediction as Translation between Augmented Natural Languages
Giovanni Paolini, Ben Athiwaratkun, Jason Krone, Alessandro Achille, Rishita Anubhai, Cícero Nogueira dos Santos, Bing Xiang, Stefano Soatto
ICLR7
2021 Mixed-Curvature Multi-Relational Graph Neural Network for Knowledge Graph Completion
abstract
Knowledge graphs (KGs) have gradually become valuable assets for many AI applications. In a KG, a node denotes an entity, and an edge (or link) denotes a relationship between the entities represented by the nodes. Knowledge graph completion infers and predicts missing edges in a KG automatically. Knowledge graph embeddings have shed light on addressing this task. Recent research embeds KGs in hyperbolic (negatively curved) space instead of conventional Euclidean (zero curved) space and is effective in capturing hierarchical structures. However, as multi-relational graphs, KGs are not structured uniformly and display intrinsic heterogeneous structures. They usually contain rich types of structures, such as hierarchical and cyclic typed structures. Embedding KGs in single-curvature space, such as Euclidean or hyperbolic space, overlooks the intrinsic heterogeneous structures of KGs, and therefore cannot accurately capture their structures. To address this issue, we propose Mixed-Curvature Multi-Relational Graph Neural Network (M2GNN), a generic approach that embeds multi-relational KGs in a mixed-curvature space for knowledge graph completion. Specifically, we define and construct a mixed-curvature space through a product manifold combining multiple single-curvature spaces (e.g., spherical, hyperbolic, or Euclidean) with the purpose of modeling a variety of structures. However, constructing a mixed-curvature space typically requires manually defining the fixed curvatures, which needs domain knowledge and additional data analysis. Improperly defined curvature space also cannot capture the structures of KGs accurately. To address this problem, we set mixed-curvatures as trainable parameters to better capture the underlying structures of the KGs. Furthermore, we propose a Graph Neural Updater by leveraging the heterogeneous relational context in mixed-curvature space to improve the quality of the embedding. Experiments on three KG datasets demonstrate that the proposed M2GNN can outperform its single geometry counterpart as well as state-of-the-art embedding methods on the KG completion task.
Shen Wang 0005, Xiaokai Wei, Cícero Nogueira dos Santos, Zhiguo Wang 0006, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang, Philip S. Yu, Isabel F. Cruz
WWW3
2020 Learning Implicit Text Generation via Feature Matching
abstract
Inkit Padhi, Pierre Dognin, Ke Bai, Cícero Nogueira dos Santos, Vijil Chenthamarakshan, Youssef Mroueh, Payel Das. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Inkit Padhi, Pierre L. Dognin, Cícero Nogueira dos Santos, Vijil Chenthamarakshan, Youssef Mroueh
ACL4
2020 Augmented Natural Language for Generative Sequence Labeling
abstract
We propose a generative framework for joint sequence labeling and sentence-level classification.Our model performs multiple sequence labeling tasks at once using a single, shared natural language output space.Unlike prior discriminative methods, our model naturally incorporates label semantics and shares knowledge across tasks.Our framework is general purpose, performing well on fewshot, low-resource, and high-resource tasks.We demonstrate these advantages on popular named entity recognition, slot labeling, and intent classification benchmarks.We set a new state-of-the-art for few-shot slot labeling, improving substantially upon the previous 5-shot (75.0%!90.9%) and 1-shot (70.4% !81.0%) state-of-the-art results.Furthermore, our model generates large improvements (46.27% !63.83%) in low-resource slot labeling over a BERT baseline by incorporating label semantics.We also maintain competitive results on high-resource tasks, performing within two points of the state-of-theart on all tasks and setting a new state-of-theart on the SNIPS dataset.
Ben Athiwaratkun, Cícero Nogueira dos Santos, Jason Krone, Bing Xiang
EMNLP (1)2
2020 DualTKB: A Dual Learning Bridge between Text and Knowledge Base
abstract
In this work, we present a dual learning approach for unsupervised text to path and path to text transfers in Commonsense Knowledge Bases (KBs).We investigate the impact of weak supervision by creating a weakly supervised dataset and show that even a slight amount of supervision can significantly improve the model performance and enable better-quality transfers.We examine different model architectures, and evaluation metrics, proposing a novel Commonsense KB completion metric tailored for generative models.Extensive experimental results show that the proposed method compares very favorably to the existing baselines.This approach is a viable step towards a more advanced system for automatic KB construction/expansion and the reverse operation of KB conversion to coherent textual descriptions.
Pierre L. Dognin, Igor Melnyk, Inkit Padhi, Cícero Nogueira dos Santos
EMNLP (1)4
2020 Beyond [CLS] through Ranking by Generation
abstract
Generative models for Information Retrieval, where ranking of documents is viewed as the task of generating a query from a document's language model, were very successful in various IR tasks in the past.However, with the advent of modern deep neural networks, attention has shifted to discriminative ranking functions that model the semantic similarity of documents and queries instead.Recently, deep generative models such as GPT2 and BART have been shown to be excellent text generators, but their effectiveness as rankers have not been demonstrated yet.In this work, we revisit the generative framework for information retrieval and show that our generative approaches are as effective as state-of-the-art semantic similarity-based discriminative models for the answer selection task.Additionally, we demonstrate the effectiveness of unlikelihood losses for IR.
Cícero Nogueira dos Santos, Xiaofei Ma 0001, Ramesh Nallapati, Zhiheng Huang, Bing Xiang
EMNLP (1)1
2020 End-to-End Synthetic Data Generation for Domain Adaptation of Question Answering Systems
abstract
Siamak Shakeri, Cicero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Feng Nan, Zhiguo Wang, Ramesh Nallapati, Bing Xiang. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Siamak Shakeri, Cícero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Feng Nan, Zhiguo Wang 0006, Ramesh Nallapati, Bing Xiang
EMNLP (1)2
2020 H2KGAT: Hierarchical Hyperbolic Knowledge Graph Attention Network
Shen Wang 0005, Xiaokai Wei, Cícero Nogueira dos Santos, Zhiguo Wang 0006, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang, Philip S. Yu
EMNLP (1)3
2019 Learning Implicit Generative Models by Matching Perceptual Features
abstract
Perceptual features (PFs) have been used with great success in tasks such as transfer learning, style transfer, and super-resolution. However, the efficacy of PFs as key source of information for learning generative models is not well studied. We investigate here the use of PFs in the context of learning implicit generative models through moment matching (MM). More specifically, we propose a new effective MM approach that learns implicit generative models by performing mean and covariance matching of features extracted from pretrained ConvNets. Our proposed approach improves upon existing MM methods by: (1) breaking away from the problematic min/max game of adversarial learning; (2) avoiding online learning of kernel functions; and (3) being efficient with respect to both number of used moments and required minibatch size. Our experimental results demonstrate that, due to the expressiveness of PFs from pretrained deep ConvNets, our method achieves state-of-the-art results for challenging benchmarks.
Cícero Nogueira dos Santos, Youssef Mroueh, Inkit Padhi, Pierre L. Dognin
ICCV1
2019 Wasserstein Barycenter Model Ensembling
Pierre L. Dognin, Igor Melnyk, Youssef Mroueh, Jerret Ross, Cícero Nogueira dos Santos, Tom Sercu
ICLR (Poster)5
2019 Sobolev Independence Criterion
abstract
We propose the Sobolev Independence Criterion (SIC), an interpretable dependency measure between a high dimensional random variable X and a response variable Y. SIC decomposes to the sum of feature importance scores and hence can be used for nonlinear feature selection. SIC can be seen as a gradient regularized Integral Probability Metric (IPM) between the joint distribution of the two random variables and the product of their marginals. We use sparsity inducing gradient penalties to promote input sparsity of the critic of the IPM. In the kernel version we show that SIC can be cast as a convex optimization problem by introducing auxiliary variables that play an important role in feature selection as they are normalized feature importance scores. We then present a neural version of SIC where the critic is parameterized as a homogeneous neural network, improving its representation power as well as its interpretability. We conduct experiments validating SIC for feature selection in synthetic and real-world experiments. We show that SIC enables reliable and interpretable discoveries, when used in conjunction with the holdout randomization test and knockoffs to control the False Discovery Rate. Code is available at http://github.com/ibm/sic.
Youssef Mroueh, Tom Sercu, Mattia Rigotti, Inkit Padhi, Cícero Nogueira dos Santos
NeurIPS5
2017 Improved Neural Relation Detection for Knowledge Base Question Answering
abstract
Relation detection is a core component of many NLP applications including Knowledge Base Question Answering (KBQA).In this paper, we propose a hierarchical recurrent neural network enhanced by residual learning which detects KB relations given an input question.Our method uses deep residual bidirectional LSTMs to compare questions and relation names via different levels of abstraction.Additionally, we propose a simple KBQA system that integrates entity linking and our proposed relation detector to make the two components enhance each other.Our experimental results show that our approach not only achieves outstanding relation detection performance, but more importantly, it helps our KBQA system achieve state-of-the-art accuracy for both single-relation (SimpleQuestions) and multi-relation (WebQSP) QA benchmarks.
Mo Yu, Wenpeng Yin 0001, Kazi Saidul Hasan, Cícero Nogueira dos Santos, Bing Xiang, Bowen Zhou 0002
ACL (1)4
2017 A Structured Self-Attentive Sentence Embedding
Zhouhan Lin, Minwei Feng, Cícero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou 0002, Yoshua Bengio
ICLR (Poster)3
2017 Domain adaptation of POS taggers without handcrafted features
abstract
Unsupervised domain adaptation is an attractive option when labeled data is lacking for some domain of interest but is available for other domain. Part-of-speech (POS) tagging is often considered a solved task when enough labeled data is available in the domain of interest. However, when considering a domain adaptation scenario, this is far from true. Several approaches have been proposed for domain adaptation of POS taggers, however as far as we know, all of them are based on handcrafted features. In this work, we employ a machine learning method whose input is exclusively composed of the raw text. This method learns word- and character-level representations (embeddings), and has been successfully applied to intra-domain tasks. We show that this method achieves strong performances on the domain adaptation of English and Portuguese POS taggers.
Irving Muller Rodrigues, Eraldo Rezende Fernandes, Cícero Nogueira dos Santos
IJCNN3
2016 Improved Representation Learning for Question Answer Matching
abstract
Passage-level question answer matching is a challenging task since it requires effective representations that capture the complex semantic relations between questions and answers.In this work, we propose a series of deep learning models to address passage answer selection.To match passage answers to questions accommodating their complex semantic relations, unlike most previous work that utilizes a single deep learning structure, we develop hybrid models that process the text using both convolutional and recurrent neural networks, combining the merits on extracting linguistic information from both structures.Additionally, we also develop a simple but effective attention mechanism for the purpose of constructing better answer representations according to the input question, which is imperative for better modeling long answer sequences.The results on two public benchmark datasets, InsuranceQA and TREC-QA, show that our proposed models outperform a variety of strong baselines.
Cícero Nogueira dos Santos, Bing Xiang, Bowen Zhou 0006
ACL (1)2
2016 Abstractive Text Summarization using Sequence-to-sequence RNNs and Beyond
abstract
In this work, we model abstractive text summarization using Attentional Encoder-Decoder Recurrent Neural Networks, and show that they achieve state-of-the-art performance on two different corpora.We propose several novel models that address critical problems in summarization that are not adequately modeled by the basic architecture, such as modeling key-words, capturing the hierarchy of sentence-toword structure, and emitting words that are rare or unseen at training time.Our work shows that many of our proposed models contribute to further improvement in performance.We also propose a new dataset consisting of multi-sentence summaries, and establish performance benchmarks for further research.
Ramesh Nallapati, Bowen Zhou 0002, Cícero Nogueira dos Santos, Caglar Gulcehre, Bing Xiang
CoNLL3
2015 Classifying Relations by Ranking with Convolutional Neural Networks
abstract
Cícero dos Santos, Bing Xiang, Bowen Zhou. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Cícero Nogueira dos Santos, Bing Xiang, Bowen Zhou 0006
ACL (1)1
2015 Detecting Semantically Equivalent Questions in Online User Forums
abstract
Two questions asking the same thing could be too different in terms of vocabulary and syntactic structure, which makes identifying their semantic equivalence challenging. This study aims to detect semantically equivalent questions in online user forums. We perform an extensive number of experiments using data from two different Stack Exchange forums. We compare standard machine learning methods such as Support Vector Machines (SVM) with a convolutional neural network (CNN). The proposed CNN generates distributed vector representations for pairs of questions and scores them using a similarity metric. We evaluate in-domain word embeddings versus the ones trained with Wikipedia, estimate the impact of the training set size, and evaluate some aspects of domain adaptation. Our experimental results show that the convolutional neural network with in-domain word embeddings achieves high performance even with limited training data.
Dasha Bogdanova, Cícero Nogueira dos Santos, Luciano Barbosa, Bianca Zadrozny
CoNLL2
2014 Deep Convolutional Neural Networks for Sentiment Analysis of Short Texts
Cícero Nogueira dos Santos, Maíra Gatti de Bayser
COLING1
2014 Learning Character-level Representations for Part-of-Speech Tagging
abstract
Distributed word representations have recently been proven to be an invaluable resource for NLP. These representations are normally learned using neural networks and capture syntactic and semantic information about words. Information about word morphology and shape is normally ignored when learning word representations. However, for tasks like part-of-speech tagging, intra-word information is extremely useful, specially when dealing with morphologically rich languages. In this paper, we propose a deep neural network that learns character-level representation of words and associate them with usual word representations to perform POS tagging. Using the proposed approach, while avoiding the use of any handcrafted feature, we produce state-of-the-art POS taggers for two languages: English, with 97.32% accuracy on the Penn Treebank WSJ corpus; and Portuguese, with 97.47% accuracy on the Mac-Morpho corpus, where the latter represents an error reduction of 12.2% on the best previous known result.
Cícero Nogueira dos Santos, Bianca Zadrozny
ICML1
2014 Latent Trees for Coreference Resolution
abstract
We describe a structure learning system for unrestricted coreference resolution that explores two key modeling techniques: latent coreference trees and automatic entropy-guided feature induction. The latent tree modeling makes the learning problem computationally feasible because it incorporates a meaningful hidden structure. Additionally, using an automatic feature induction method, we can efficiently build enhanced nonlinear models using linear model learning algorithms. We present empirical results that highlight the contribution of each modeling technique used in the proposed system. Empirical evaluation is performed on the multilingual unrestricted coreference CoNLL-2012 Shared Task datasets, which comprise three languages: Arabic, Chinese and English. We apply the same system to all languages, except for minor adaptations to some language-dependent features such as nested mentions and specific static pronoun lists. A previous version of this system was submitted to the CoNLL-2012 Shared Task closed track, achieving an official score of 58.69, the best among the competitors. The unique enhancement added to the current system version is the inclusion of candidate arcs linking nested mentions for the Chinese language. By including such arcs, the score increases by almost 4.5 points for that language. The current system shows a score of 60.15, which corresponds to a 3.5% error reduction, and is the best performing system for each of the three languages.
Eraldo Rezende Fernandes, Cícero Nogueira dos Santos, Ruy Milidiú
Comput. Linguistics2
2013 Large-Scale Multi-agent-Based Modeling and Simulation of Microblogging-Based Online Social Network
Maíra Gatti de Bayser, Paulo Rodrigo Cavalin, Samuel Martins Barbosa Neto, Claudio S. Pinhanez, Cícero Nogueira dos Santos, Daniel Gribel, Ana Paula Appel
MABS5
2010 ETL Ensembles for Chunking, NER and SRL
Cícero Nogueira dos Santos, Ruy Milidiú, Carlos E. M. Crestana, Eraldo Rezende Fernandes
CICLing1
2008 Phrase Chunking Using Entropy Guided Transformation Learning
Ruy Milidiú, Cícero Nogueira dos Santos, Julio C. Duarte
ACL2
2007 Probabilistic Classifications with TBL
Cícero Nogueira dos Santos, Ruy Milidiú
CICLing1