TY - GEN
T1 - Building a corpus of spatial relational expressions extracted from web documents
AU - Wallgrün, Jan Oliver
AU - Klippel, Alexander
AU - Baldwin, Timothy
N1 - Publisher Copyright:
Copyright 2014 ACM.
PY - 2014/11/4
Y1 - 2014/11/4
N2 - Spatial language, despite decades of research, still poses substantial challenges for automated systems, for instance in geographic information retrieval or human-robot interaction. We describe an approach to building a corpus of natural language expressions extracted from web documents for analyzing and modeling spatial relational expressions (SRE). The unique characteristic of this corpus is that it is built around georeferenced triplets, with each triplet containing two entities (including their latitude/longitude coordinates) related by a spatial expression such as near. While the approach is still experimental, our first results are promising, in that we believe they will form the foundation for a comprehensive contextualized model for interpreting spatial natural language expressions. For the time being, we are focusing on a single domain, hotel reviews. This domain restriction allowed us to implement a proof-of-concept that this approach, with advances in natural language technologies, will indeed deliver a comprehensive corpus. The potential to collect larger corpora, and associated challenges, is discussed.
AB - Spatial language, despite decades of research, still poses substantial challenges for automated systems, for instance in geographic information retrieval or human-robot interaction. We describe an approach to building a corpus of natural language expressions extracted from web documents for analyzing and modeling spatial relational expressions (SRE). The unique characteristic of this corpus is that it is built around georeferenced triplets, with each triplet containing two entities (including their latitude/longitude coordinates) related by a spatial expression such as near. While the approach is still experimental, our first results are promising, in that we believe they will form the foundation for a comprehensive contextualized model for interpreting spatial natural language expressions. For the time being, we are focusing on a single domain, hotel reviews. This domain restriction allowed us to implement a proof-of-concept that this approach, with advances in natural language technologies, will indeed deliver a comprehensive corpus. The potential to collect larger corpora, and associated challenges, is discussed.
UR - http://www.scopus.com/inward/record.url?scp=84942411097&partnerID=8YFLogxK
UR - http://www.scopus.com/inward/citedby.url?scp=84942411097&partnerID=8YFLogxK
U2 - 10.1145/2675354.2675702
DO - 10.1145/2675354.2675702
M3 - Conference contribution
AN - SCOPUS:84942411097
T3 - Proceedings of the 8th Workshop on Geographic Information Retrieval, GIR 2014
BT - Proceedings of the 8th Workshop on Geographic Information Retrieval, GIR 2014
A2 - Purves, Ross S.
A2 - Jones, Christopher B.
PB - Association for Computing Machinery, Inc
T2 - 8th Workshop on Geographic Information Retrieval, GIR 2014
Y2 - 4 November 2014 through 7 November 2014
ER -