Workshop

The workshop “DutchSemCor” at the VU University Amsterdam was held on Thursday 25 October 2012.

In this workshop, we present the results of the three-year NWO medium investment subsidy project DutchSemCor. The goal of the project was to deliver a one-million word Dutch corpus that is fully sense-tagged with senses and domain tags from the Cornetto database. The corpus data is based on existing corpus material collected in the projects CGN, D-CoI and SoNaR. These corpora were extended where necessary to find sufficient examples for meanings of words that are less frequent and do not appear in the above corpora. The resulting corpus is extremely rich in terms of lexical semantic information. Its availability enables many new lines of research and technology developments for the Dutch language such as the development of word sense disambiguation systems.

The organizers Piek Vossen (piek.vossen@vu.nl) & Attila Görög (a.gorog@vu.nl)

The presentations of the workshop are now online and can be downloaded here: