Building a corpus of spatial relational expressions extracted from web documents

Jan Oliver Wallgrün, Alexander Klippel, Timothy Baldwin

Research output: Chapter in Book/Report/Conference proceedingConference contribution

20 Scopus citations

Abstract

Spatial language, despite decades of research, still poses substantial challenges for automated systems, for instance in geographic information retrieval or human-robot interaction. We describe an approach to building a corpus of natural language expressions extracted from web documents for analyzing and modeling spatial relational expressions (SRE). The unique characteristic of this corpus is that it is built around georeferenced triplets, with each triplet containing two entities (including their latitude/longitude coordinates) related by a spatial expression such as near. While the approach is still experimental, our first results are promising, in that we believe they will form the foundation for a comprehensive contextualized model for interpreting spatial natural language expressions. For the time being, we are focusing on a single domain, hotel reviews. This domain restriction allowed us to implement a proof-of-concept that this approach, with advances in natural language technologies, will indeed deliver a comprehensive corpus. The potential to collect larger corpora, and associated challenges, is discussed.

Original languageEnglish (US)
Title of host publicationProceedings of the 8th Workshop on Geographic Information Retrieval, GIR 2014
EditorsRoss S. Purves, Christopher B. Jones
PublisherAssociation for Computing Machinery, Inc
ISBN (Electronic)9781450331357
DOIs
StatePublished - Nov 4 2014
Event8th Workshop on Geographic Information Retrieval, GIR 2014 - Dallas, United States
Duration: Nov 4 2014Nov 7 2014

Publication series

NameProceedings of the 8th Workshop on Geographic Information Retrieval, GIR 2014

Other

Other8th Workshop on Geographic Information Retrieval, GIR 2014
Country/TerritoryUnited States
CityDallas
Period11/4/1411/7/14

All Science Journal Classification (ASJC) codes

  • Geography, Planning and Development

Fingerprint

Dive into the research topics of 'Building a corpus of spatial relational expressions extracted from web documents'. Together they form a unique fingerprint.

Cite this