VLDB 2026 Research / reviewers in the wild / expert
Freda Shi
dblp:194/2512 · also Haoyue Shi 0001
· DBLP profile ↗
27ranked-venue papers
13as first author
21since 2021 · last 2026
0009-0009-5697-449XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 12 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | How Tokenization Limits Phonological Knowledge Representation in Language Models and How to Improve ThemabstractTokenization is the first step in every language model (LM), yet it never takes the sounds of words into account.We investigate how tokenization influences text-only LMs' ability to represent phonological knowledge.Through a series of probing experiments, we show that subword-based tokenization systematically weakens the encoding of both local (e.g., rhyme) and global (e.g., syllabification) phonological features.To quantify this effect, we introduce the syllabification-tokenization alignment distance (STAD), a metric that measures the misalignment between a model's tokenization and the natural syllable boundaries of words, and find that higher misalignment correlates with poorer phonological representations, providing a simple diagnostic for phonologyaware tokenization.To address these limitations, we propose a lightweight IPA-based finetuning method that infuses phonological awareness into LMs, leading to consistent improvements across three phonology-related tasks while largely preserving math and general reasoning ability, with 1.1% and 0.9% drops on GSM8K and MMLU, respectively.1 Disen Liao, Freda Shi |
ACL (1) | 2 |
| 2026 | On the Effect of Hyperparameters in Language Modeling for Computational LinguisticsabstractTraining language models and examining their linguistic behaviors have been a common protocol in computational linguistics for studying linguistic phenomena and modeling human language processing.However, work in this area is often limited to proof-of-concept demonstrations with arbitrary model configurations, without considering hyperparameter sensitivity, an important source of variation in model performance.In this work, we replicate three prior studies (Chang and Bergen, 2022; Hu et al., 2020b;Kuribayashi et al., 2024) with hyperparameters varied within a practical range, and show that modest hyperparameter changes can alter some qualitative conclusions about models' linguistic abilities and even reverse the ranking of model performance.Our results highlight the risk that prior work may have reflected optimization artifacts rather than the genuine inductive biases of model classes, and that hyperparameter sensitivity should receive more attention as a factor that can meaningfully influence model behavior.We suggest future work to report the variation of performance across the configuration space to enhance the reliability and generalizability of conclusions. Ruoxi Ning, Yongpeng Zhu, Qingcheng Zeng, Tatsuki Kuribayashi, Freda Shi |
ACL (1) | 5 |
| 2026 | DriveLegal: Toward legally compliant driving via trustworthy hybrid retrieval-augmented LLMsabstract• Modular legal-interpretation layer with hybrid vector–graph RAG for AV guidance. • Two datasets: SFT and RAG for multilingual, cross-jurisdiction evaluation. • Hybrid retrieval improves faithfulness and reduces hallucination vs single modes. • Trust module scores context, groundedness, and answer relevance online. • Validated in smart-cabin, V2X intersection monitoring, and offline auditing. Autonomous vehicles (AVs) face persistent challenges in complying with complex and evolving traffic laws. Existing approaches, including rule-based, learning-based, and large language model (LLM) methods, each face limits in adaptability, generalizability, or trustworthiness. We present DriveLegal , a modular legal-interpretation framework for downstream autonomous driving applications. DriveLegal pairs fine-tuned multilingual large language models (LLMs) with an intelligent hybrid retrieval module that routes between vector search and knowledge graph, then returns concise, cited answers. A trust layer scores context relevance, groundedness, and answer relevance and supports continuous improvement through periodic automatic signals and targeted human review. We introduce the DriveLegal datasets for supervised fine-tuning and for retrieval and graph reasoning. Across benchmarks and case studies in smart cabin and vehicle-to-everything (V2X) settings, the hybrid retrieval strategy improves contextual accuracy and reduces hallucination while producing jurisdiction-aware outputs suitable for compliance checks, incident analysis, and reporting. Shucheng Huang, Chen Sun 0008, Minghao Ning, Changye Ma, Jiaming Zhong, Keqi Shu, Freda Shi, Amir Khajepour |
Expert Syst. Appl. | 8 |
| 2025 | Learning Language Structures Through GroundingabstractLanguage is highly structured, with syntactic and semantic structures, to some extent, agreed upon by speakers. With implicit or explicit awareness of such structures, humans can learn and use language efficiently and generalize to sentences that contain unseen words. Motivated by human language learning, in this presentation, I will introduce a family of machine learning tasks that learns language structures through grounding, where distant supervision from other data sources (i.e., grounds), including but not limited to different modalities (e.g., vision), execution results of programs, and other languages, are used to guide the learning of language structures. I will demonstrate the potential of this task formulation, advocate for its adoption through three schemes, and discuss the possibility of the general language learning problem through grounding. Freda Shi |
AAAI | 1 |
| 2025 | SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic DataabstractVision-language models (VLMs) work well in tasks ranging from image captioning to visual question answering (VQA), yet they struggle with spatial reasoning, a key skill for understanding our physical world that humans excel at.We find that spatial relations are generally rare in widely used VL datasets, with only a few being well represented while most form a long tail of underrepresented relations.This gap leaves VLMs ill-equipped to handle diverse spatial relationships.To bridge it, we construct a synthetic VQA dataset focused on spatial reasoning generated from hyperdetailed image descriptions in Localized Narratives, DOCCI, and PixMo-Cap.Our dataset consists of 455k samples containing 3.4 million QA pairs.Trained on this dataset, our Spatial-Reasoning Enhanced (SpaRE) VLMs show strong improvements on spatial reasoning benchmarks, achieving up to a 49% performance gain on the What's Up benchmark, while maintaining strong results on general tasks.Our work narrows the gap between human and VLM spatial reasoning and makes VLMs more capable in real-world tasks such as robotics and navigation.We plan to share our code and dataset in due course. Michael Ogezi, Freda Shi |
ACL (1) | 2 |
| 2025 | Logical forms complement probability in understanding language model (and human) performanceabstractWith the increasing interest in using large language models (LLMs) for planning in natural language, understanding their behaviors becomes an important research question.This work conducts a systematic investigation of LLMs' ability to perform logical reasoning in natural language.We introduce a controlled dataset of hypothetical and disjunctive syllogisms in propositional and modal logic and use it as the testbed for understanding LLM performance.Our results lead to novel insights in predicting LLM behaviors: in addition to the probability of input (Gonen et al., 2023;McCoy et al., 2024), logical forms should be considered as important factors.In addition, we show similarities and discrepancies between the logical reasoning performances of humans and LLMs by collecting and comparing behavioral data from both. Freda Shi |
ACL (1) | 2 |
| 2025 | Distribution Prompting: Understanding the Expressivity of Language Models Through the Next-Token Distributions They Can ProduceabstractAutoregressive neural language models (LMs) generate a probability distribution over tokens at each time step given a prompt.In this work, we attempt to systematically understand the probability distributions that LMs can produce, showing that some distributions are significantly harder to elicit than others.Specifically, for any target next-token distribution over the vocabulary, we attempt to find a prompt that induces the LM to output a distribution as close as possible to the target, using either soft (Li and Liang, 2021) or hard (Wallace et al., 2019) gradient-based prompt tuning.We find that (1) in general, distributions with very low or very high entropy are easier to approximate than those with moderate entropy; (2) among distributions with the same entropy, those containing "outlier tokens" are easier to approximate; (3) target distributions generated by LMs-even LMs with different tokenizers-are easier to approximate than randomly chosen targets.These results offer insights into the expressiveness of LMs and the challenges of using them as probability distribution proposers. Haojin Wang, Zining Zhu 0001, Freda Shi |
EMNLP | 3 |
| 2025 | LingGym: How Far Are LLMs from Thinking Like Field Linguists?abstractThis paper introduces LINGGYM, a new benchmark that evaluates LLMs' capacity for metalinguistic reasoning using Interlinear Glossed Text (IGT) and grammatical descriptions extracted from 18 typologically diverse reference grammars.Unlike previous work that focuses on specific downstream tasks, we assess whether LLMs can generalize linguistic inference across low-resource languages and structures not seen during training.We present a controlled evaluation task: Word-Gloss Inference, in which the model must infer a missing word and gloss from context using varying levels of linguistic information (e.g., glosses, grammatical explanations, translations).Our results show that incorporating structured linguistic cues leads to consistent improvements in reasoning performance across all models.This work highlights both the promise and current limitations of using LLMs for typologically informed linguistic analysis and low-resource language documentation. Changbing Yang, Franklin Ma, Freda Shi |
EMNLP | 3 |
| 2025 | Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference under AmbiguitiesabstractSpatial expressions in situated communication can be ambiguous, as their meanings vary depending on the frames of reference (FoR) adopted by speakers and listeners. While spatial language understanding and reasoning by vision-language models (VLMs) have gained increasing attention, potential ambiguities in these models are still under-explored. To address this issue, we present the COnsistent Multilingual Frame Of Reference Test (COMFORT), an evaluation protocol to systematically assess the spatial reasoning capabilities of VLMs. We evaluate nine state-of-the-art VLMs using COMFORT. Despite showing some alignment with English conventions in resolving ambiguities, our experiments reveal significant shortcomings of VLMs: notably, the models (1) exhibit poor robustness and consistency, (2) lack the flexibility to accommodate multiple FoRs, and (3) fail to adhere to language-specific or culture-specific conventions in cross-lingual tests, as English tends to dominate other languages. With a growing effort to align vision-language models with human cognitive intuitions, we call for more attention to the ambiguous nature and cross-cultural diversity of spatial reasoning. Fengyuan Hu, Jayjun Lee, Freda Shi, Parisa Kordjamshidi, Joyce Y. Chai, Ziqiao Ma 0001 |
ICLR | 4 |
| 2024 | LogogramNLP: Comparing Visual and Textual Representations of Ancient Logographic Writing Systems for NLPabstractStandard natural language processing (NLP) pipelines operate on symbolic representations of language, which typically consist of sequences of discrete tokens.However, creating an analogous representation for ancient logographic writing systems is an extremely laborintensive process that requires expert knowledge.At present, a large portion of logographic data persists in a purely visual form due to the absence of transcription-this issue poses a bottleneck for researchers seeking to apply NLP toolkits to study ancient logographic languages: most of the relevant data are images of writing.This paper investigates whether direct processing of visual representations of language offers a potential solution.We introduce LogogramNLP, the first benchmark enabling NLP analysis of ancient logographic languages, featuring both transcribed and visual datasets for four writing systems along with annotations for tasks like classification, translation, and parsing.Our experiments compare systems that employ recent visual and text encoding strategies as backbones.The results demonstrate that visual representations outperform textual representations for some investigated tasks, suggesting that visual processing pipelines may unlock a large amount of cultural heritage data of logographic languages for NLP-based analyses. Danlu Chen, Freda Shi, Aditi Agarwal, Jacobo Myerston, Taylor Berg-Kirkpatrick |
ACL (1) | 2 |
| 2024 | Structured Tree Alignment for Evaluation of (Speech) Constituency ParsingabstractWe present the structured average intersectionover-union ratio (STRUCT-IOU), a similarity metric between constituency parse trees motivated by the problem of evaluating speech parsers.STRUCT-IOU enables comparison between a constituency parse tree (over automatically recognized spoken word boundaries) with the ground-truth parse (over written words).To compute the metric, we project the groundtruth parse tree to the speech domain by forced alignment, align the projected ground-truth constituents with the predicted ones under certain structured constraints, and calculate the average IOU score across all aligned constituent pairs.STRUCT-IOU takes word boundaries into account and overcomes the challenge that the predicted words and ground truth may not have perfect one-to-one correspondence.Extending to the evaluation of text constituency parsing, we demonstrate that STRUCT-IOU can address token-mismatch issues, and shows higher tolerance to syntactically plausible parses than PARSEVAL (Black et al., 1991). 1 Freda Shi, Kevin Gimpel, Karen Livescu |
ACL (1) | 1 |
| 2024 | Gated Slot Attention for Efficient Linear-Time Sequence ModelingabstractLinear attention Transformers and their gated variants, celebrated for enabling parallel training and efficient recurrent inference, still fall short in recall-intensive tasks compared to traditional Transformers and demand significant resources for training from scratch.
This paper introduces Gated Slot Attention (GSA), which enhances Attention with Bounded-memory-Control (ABC) by incorporating a gating mechanism inspired by Gated Linear Attention (GLA).
Essentially, GSA comprises a two-layer GLA linked via $\operatorname{softmax}$, utilizing context-aware memory reading and adaptive forgetting to improve memory capacity while maintaining compact recurrent state size.
This design greatly enhances both training and inference efficiency through GLA's hardware-efficient training algorithm and reduced state size.
Additionally, retaining the $\operatorname{softmax}$ operation is particularly beneficial in ``finetuning pretrained Transformers to RNNs'' (T2R) settings, reducing the need for extensive training from scratch.
Extensive experiments confirm GSA's superior performance in scenarios requiring in-context recall and in T2R settings. Yu Zhang 0092, Rui-Jie Zhu 0003, Yue Zhang 0004, Leyang Cui, Yiqiao Wang 0005, Bolun Wang, Freda Shi, Bailin Wang, Wei Bi, Peng Zhou 0017, Guohong Fu |
NeurIPS | 8 |
| 2023 | Audio-Visual Neural Syntax AcquisitionabstractWe study phrase structure induction from visually-grounded speech. The core idea is to first segment the speech waveform into sequences of word segments, and subsequently induce phrase structure using the inferred segment-level continuous representations. We present the Audio-Visual Neural Syntax Learner (AV-NSL) that learns phrase structure by listening to audio and looking at images, without ever being exposed to text. By training on paired images and spoken captions, AV-NSL exhibits the capability to infer meaningful phrase structures that are comparable to those derived by naturally-supervised text parsers, for both English and German. Our findings extend prior work in unsupervised language acquisition from speech and grounded grammar induction, and present one approach to bridge the gap between the two topics. Cheng-I Lai, Freda Shi, Puyuan Peng, Kevin Gimpel, Shiyu Chang, Yung-Sung Chuang, Saurabhchand Bhati, David D. Cox, David F. Harwath, Yang Zhang 0001, Karen Livescu, James R. Glass |
ASRU | 2 |
| 2023 | InCoder: A Generative Model for Code Infilling and Synthesis
Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida I. Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Scott Yih, Luke Zettlemoyer, Mike Lewis |
ICLR | 6 |
| 2023 | Language models are multilingual chain-of-thought reasoners
Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang 0002, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das 0001, Jason Wei |
ICLR | 1 |
| 2023 | Large Language Models Can Be Easily Distracted by Irrelevant ContextabstractLarge language models have achieved impressive performance on various natural language processing tasks. However, so far they have been evaluated primarily on benchmarks where all information in the input context is relevant for solving the task. In this work, we investigate the *distractibility* of large language models, i.e., how the model prediction can be distracted by irrelevant context. In particular, we introduce Grade-School Math with Irrelevant Context (GSM-IC), an arithmetic reasoning dataset with irrelevant information in the problem description. We use this benchmark to measure the distractibility of different prompting techniques for large language models, and find that the model is easily distracted by irrelevant information. We also identify several approaches for mitigating this deficiency, such as decoding with self-consistency and adding to the prompt an instruction that tells the language model to ignore the irrelevant information. Freda Shi, Kanishka Misra, Nathan Scales, David Dohan, Ed H. Chi, Nathanael Schärli, Denny Zhou |
ICML | 1 |
| 2022 | Deep Clustering of Text Representations for Supervision-Free Probing of SyntaxabstractWe explore deep clustering of multilingual text representations for unsupervised model interpretation and induction of syntax. As these representations are high-dimensional, out-of-the-box methods like K-means do not work well. Thus, our approach jointly transforms the representations into a lower-dimensional cluster-friendly space and clusters them. We consider two notions of syntax: Part of Speech Induction (POSI) and Constituency Labelling (CoLab) in this work. Interestingly, we find that Multilingual BERT (mBERT) contains surprising amount of syntactic knowledge of English; possibly even as much as English BERT (E-BERT). Our model can be used as a supervision-free probe which is arguably a less-biased way of probing. We find that unsupervised probes show benefits from higher layers as compared to supervised probes. We further note that our unsupervised probe utilizes E-BERT and mBERT representations differently, especially for POSI. We validate the efficacy of our probe by demonstrating its capabilities as a unsupervised syntax induction technique. Our probe works well for both syntactic formalisms by simply adapting the input representations. We report competitive performance of our probe on 45-tag English POSI, state-of-the-art performance on 12-tag POSI across 10 languages, and competitive results on CoLab. We also perform zero-shot syntax induction on resource impoverished languages and report strong results. Vikram Gupta, Freda Shi, Kevin Gimpel, Mrinmaya Sachan |
AAAI | 2 |
| 2022 | Substructure Distribution Projection for Zero-Shot Cross-Lingual Dependency ParsingabstractWe present substructure distribution projection (SUBDP), a technique that projects a distribution over structures in one domain to another, by projecting substructure distributions separately.Models for the target domain can then be trained, using the projected distributions as soft silver labels.We evaluate SUBDP on zeroshot cross-lingual dependency parsing, taking dependency arcs as substructures: we project the predicted dependency arc distributions in the source language(s) to target language(s), and train a target language parser on the resulting distributions.Given an English treebank as the only source of human supervision, SUBDP achieves better unlabeled attachment score than all prior work on the Universal Dependencies v2.2 (Nivre et al., 2020) test set across eight diverse target languages, as well as the best labeled attachment score on six languages.In addition, SUBDP improves zeroshot cross-lingual dependency parsing with very few (e.g., 50) supervised bitext pairs, across a broader range of target languages. Freda Shi, Kevin Gimpel, Karen Livescu |
ACL (1) | 1 |
| 2022 | Natural Language to Code Translation with ExecutionabstractGenerative models of code, pretrained on large corpora of programs, have shown great success in translating natural language to code (Chen et al., 2021;Austin et al., 2021; Li et al., 2022, inter alia).While these models do not explicitly incorporate program semantics (i.e., execution results) during training, they are able to generate correct solutions for many problems.However, choosing a single correct program from a generated set for each problem remains challenging.In this work, we introduce execution resultbased minimum Bayes risk decoding (MBR-EXEC) for program selection and show that it improves the few-shot performance of pretrained code models on natural-language-tocode tasks.We select output programs from a generated candidate set by marginalizing over program implementations that share the same semantics.Because exact equivalence is intractable, we execute each program on a small number of test inputs to approximate semantic equivalence.Across datasets, execution or simulated execution significantly outperforms the methods that do not involve program semantics.We find that MBR-EXEC consistently improves over all execution-unaware selection methods, suggesting it as an effective approach for natural language to code translation.1 Freda Shi, Daniel Fried, Marjan Ghazvininejad, Luke Zettlemoyer, Sida I. Wang |
EMNLP | 1 |
| 2021 | Bilingual Lexicon Induction via Unsupervised Bitext Construction and Word AlignmentabstractHaoyue Shi, Luke Zettlemoyer, Sida I. Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Freda Shi, Luke Zettlemoyer, Sida I. Wang |
ACL/IJCNLP (1) | 1 |
| 2021 | Grammar-Based Grounded Lexicon LearningabstractWe present Grammar-Based Grounded Language Learning (G2L2), a lexicalist approach toward learning a compositional and grounded meaning representation of language from grounded data, such as paired images and texts. At the core of G2L2 is a collection of lexicon entries, which map each word to a tuple of a syntactic type and a neuro-symbolic semantic program. For example, the word shiny has a syntactic type of adjective; its neuro-symbolic semantic program has the symbolic form $\lambda x.\textit{filter}(x, \textbf{SHINY})$, where the concept SHINY is associated with a neural network embedding, which will be used to classify shiny objects. Given an input sentence, G2L2 first looks up the lexicon entries associated with each token. It then derives the meaning of the sentence as an executable neuro-symbolic program by composing lexical meanings based on syntax. The recovered meaning programs can be executed on grounded inputs. To facilitate learning in an exponentially-growing compositional space, we introduce a joint parsing and expected execution algorithm, which does local marginalization over derivations to reduce the training time. We evaluate G2L2 on two domains: visual reasoning and language-driven navigation. Results show that G2L2 can generalize from small amounts of data to novel compositions of words. Jiayuan Mao, Freda Shi, Jiajun Wu 0001, Roger Levy, Josh Tenenbaum |
NeurIPS | 2 |
| 2020 | On the Role of Supervision in Unsupervised Constituency ParsingabstractWe analyze several recent unsupervised constituency parsing models, which are tuned with respect to the parsing F 1 score on the Wall Street Journal (WSJ) development set (1,700 sentences).We introduce strong baselines for them, by training an existing supervised parsing model (Kitaev and Klein, 2018) on the same labeled examples they access.When training on the 1,700 examples, or even when using only 50 examples for training and 5 for development, such a few-shot parsing approach can outperform all the unsupervised parsing methods by a significant margin.Fewshot parsing can be further improved by a simple data augmentation method and selftraining.This suggests that, in order to arrive at fair conclusions, we should carefully consider the amount of labeled data used for model development.We propose two protocols for future work on unsupervised parsing: (i) use fully unsupervised criteria for hyperparameter tuning and model selection; (ii) use as few labeled examples as possible for model development, and compare to few-shot parsing trained on the same labeled examples.1 Freda Shi, Karen Livescu, Kevin Gimpel |
EMNLP (1) | 1 |
| 2019 | Visually Grounded Neural Syntax AcquisitionabstractWe present the Visually Grounded Neural Syntax Learner (VG-NSL), an approach for learning syntactic representations and structures without explicit supervision.The model learns by looking at natural images and reading paired captions.VG-NSL generates constituency parse trees of texts, recursively composes representations for constituents, and matches them with images.We define the concreteness of constituents by their matching scores with images, and use it to guide the parsing of text.Experiments on the MSCOCO data set show that VG-NSL outperforms various unsupervised parsing approaches that do not use visual grounding, in terms of F 1 scores against gold parse trees.We find that VG-NSL is much more stable with respect to the choice of random initialization and the amount of training data.We also find that the concreteness acquired by VG-NSL correlates well with a similar measure defined by linguists.Finally, we also apply VG-NSL to multiple languages in the Multi30K data set, showing that our model consistently outperforms prior unsupervised approaches. 1 Freda Shi, Jiayuan Mao, Kevin Gimpel, Karen Livescu |
ACL (1) | 1 |
| 2018 | Learning Visually-Grounded Semantics from Contrastive Adversarial SamplesabstractWe study the problem of grounding distributional representations of texts on the visual domain, namely visual-semantic embeddings (VSE for short). Begin with an insightful adversarial attack on VSE embeddings, we show the limitation of current frameworks and image-text datasets (e.g., MS-COCO) both quantitatively and qualitatively. The large gap between the number of possible constitutions of real-world semantics and the size of parallel data, to a large extent, restricts the model to establish a strong link between textual semantics and visual concepts. We alleviate this problem by augmenting the MS-COCO image captioning datasets with textual contrastive adversarial samples. These samples are synthesized using language priors of human and the WordNet knowledge base, and enforce the model to ground learned embeddings to concrete concepts within the image. This simple but powerful technique brings a noticeable improvement over the baselines on a diverse set of downstream tasks, in addition to defending known-type adversarial attacks. Codes are available at https://github.com/ExplorerFreda/VSE-C. Freda Shi, Jiayuan Mao, Tete Xiao, Yuning Jiang 0001, Jian Sun 0001 |
COLING | 1 |
| 2018 | On Tree-Based Neural Sentence ModelingabstractNeural networks with tree-based sentence encoders have shown better results on many downstream tasks.Most of existing tree-based encoders adopt syntactic parsing trees as the explicit structure prior.To study the effectiveness of different tree structures, we replace the parsing trees with trivial trees (i.e., binary balanced tree, left-branching tree and right-branching tree) in the encoders.Though trivial trees contain no syntactic information, those encoders get competitive or even better results on all of the ten downstream tasks we investigated.This surprising result indicates that explicit syntax guidance may not be the main contributor to the superior performances of tree-based neural sentence modeling.Further analysis show that tree modeling gives better results when crucial words are closer to the final representation.Additional experiments give more clues on how to design an effective tree-based encoder.Our code is opensource and available at https://github.com/ExplorerFreda/TreeEnc. Freda Shi, Hao Zhou 0012, Jiaze Chen, Lei Li 0005 |
EMNLP | 1 |
| 2018 | Constructing High Quality Sense-specific Corpus and Word Embedding via Unsupervised Elimination of Pseudo Multi-sense
Freda Shi, Xihao Wang |
LREC | 1 |
| 2017 | Joint Saliency Estimation and Matching using Image Regions for Geo-Localization of Online VideoabstractIn this paper, we study automatic geo-localization of online event videos. Different from general image localization task through matching, the appearance of an environment during significant events varies greatly from its daily appearance, since there are usually crowds, decorations or even destruction when a major event happens. This introduces a major challenge: matching the event environment to the daily environment, e.g. as recorded by Google Street View. We observe that some regions in the image, as part of the environment, still preserve the daily appearance even though the whole image (environment) looks quite different. Based on this observation, we formulate the problem as joint saliency estimation and matching at the image region level, as opposed to the key point or whole-image level. As image-level labels of daily environment are easily generated with GPS information, we treat region based saliency estimation and matching as a weakly labeled learning problem over the training data. Our solution is to iteratively optimize saliency and the region-matching model. For saliency optimization, we derive a closed form solution, which has an intuitive explanation. For region matching model optimization, we use self-paced learning to learn from the pseudo labels generated by (sub-optimal) saliency values. We conduct extensive experiments on two challenging public datasets: Boston Marathon 2013 and Tokyo Time Machine. Experimental results show that our solution significantly improves over matching on whole images and the automatically learned saliency is a strong predictor of distinctive building areas. Freda Shi, Jia Chen 0001, Alex Hauptmann 0001 |
ICMR | 1 |