Blog
Identity of Long-Tail Entities in Text
The Meaning of Word Sense Disambiguation Research
Selene with Orange
Lenka Mask — Can machines understand human language?
Can machines understand human language?
Three central concepts
Understanding Language by Machines involves three central concepts: identity, reference and perspective.
Identity
Identity: we do not know what are all the things in the world.
Reference
Reference: a word can refer to many different things, but many different words can also refer to the same thing.
Perspective
Perspective: we can not see the world without framing.
Baratheon
There is no such thing as a perfect machine.
VU-robot Leolani in Computable Future Lab
Piek Vossen invited by ict-magazine Computable for a keynote together with Pepper robot Leolani at Infosecurity.nl, a trade fair and congress about IT security, data management and cloud computing.
Talking with Robots: Weekend of Science 2018
Talking with Robots: Children between the ages of 8 and 14 years old introduced themselves to Pepper robot Leolani, asked questions, presented objects and practiced their English, by @PiekVossen at VU workshop 06 October 2018 during Weekend of Science.
Keynote lecture 2018 PhD Day Groningen
Piek Vossen invited keynote speaker at the 7th edition of the PhD Day Groningen, Friday September 21st 2018 in the Oosterpoort in Groningen.
Interview Computable 51 #5 on Communication between Humans and Robots
Video keynote: Text, Speech and Dialogue (TSD 2018), Brno
A Reference Machine with a Theory of Mind for Social Communication
Our first Pepper robot ‘Leolani‘ publication:
To Communicate with an Imperfect Robot — Get It? At Paradiso Amsterdam
Paradiso Amsterdam, Sunday March 25, 2018 11:00 AM BUY TICKETS
— The Heavenly Voice of Leolani – Visiting Piek Vossen
Media and movies like to suggest that robots will take over our jobs, the world and even us. But is this really the case? We ask VU professor and Spinoza prize winner Piek Vossen (Faculty of Humanities) a few questions. Piek Vossen conducts research on language with his robot Leolani.
— 9th Global WordNet Conference Jan. 8—12, 2018
GWC 2018
— From Reading Machines to Reference Machines
Grounding language for machines
— Kids Talking with Robots. Workshops: Sat. Oct. 7 2017
Talking with robots
— SemEval-2018 Task: Counting Events and Participants in the Long Tail
Marten Postma, Filip Ilievski, and Piek Vossen from CLTL are hosting task 5 at SemEval-2018, entitled “Counting Events and Participants within Highly Ambiguous Data covering a very Long Tail”. This is a “referential quantification” task that requires systems to establish the meaning, reference and identity of events and participants in news articles. The task data (texts and answers) are prepared in such a way that the task deliberately exhibits large ambiguity and variation, as well as coverage of long tail phenomena by including a substantial amount of low-frequent, local events and entities. More information about this referential quantification task can be found in this handout or at the task website.
— Call for University Research Fellow
Apply for University Research Fellow 2017-2018
Deadline Friday 30 June 2017
— IJCNLP, Taipei Nov. 27—Dec. 1 2017: Semantics of the Long Tail
The workshop “SLT-1: Semantics of the Long Tail”, initiated by Piek Vossen, Filip Ilievski, and Marten Postma from VU Amsterdam, has been accepted at the next edition of the IJCNLP conference, to be held in Taipei, Taiwan on November 27 – December 1, 2017. The SLT-1 workshop is co-organized by eight external internationally recognized researchers: Eduard Hovy, Chris Welty, Martha Palmer, Ivan Titov, Philipp Cimiano, Eneko Agirre, Frank van Harmelen and Key-Sun Choi. This workshop aims at a critical discussion on the relevance and complexity of various long tail phenomena in text, i.e. hard, though non-frequent cases that need to be resolved for correct language interpretation, but are neglected by current systems. More information about the SLT-1 workshop can be found at: http://www.understandinglanguagebymachines.org/semantics-of-the-long-tail/
Language, Knowledge and People in Perspective
| About | Program | Day 1 | Day 2 | Day 3 | Day 4 | Photos | Topics | Sponsors | | — | — | — | — | — | — | — | — | — |
— Minh Le and Antske Fokkens’ paper accepted for EACL 2017
Title: Tackling Error Propagation through Reinforcement Learning: A Case of Greedy Dependency Parsing
2ND SPINOZA WORKSHOP: “LOOKING AT THE LONG TAIL”
*** Slides from the invited talks:
— In the media: Will Artificial Intelligence learn to be biased?
Last week Motherboard published an article featuring a paper by PhD student Emiel van Miltenburg. (Later also published by the Dutch Motherboard.) Van Miltenburg found that the data that is commonly used to train automatic image description systems contain stereotypes and biases, leading to the question whether computers will be biased, too.
Illustration from Stereotyping and Bias in the Flickr30k Dataset, presentation by: Emiel van Miltenburg at MMC 2016 @ LREC 2016
— An AlphaGo for Natural Language?
An AlphaGo for Natural Language?
— Semantic Class Manager
Python API to obtain semantic classes (Basic Level Concepts, WordNet Domains and WordNet SuperSenses) related to synsets in WordNet and other functionalities related with Semantic Classes. The repository can be found freely available at https://github.com/rubenIzquierdo/semantic_class_manager
— Similarity & Relatedness
Similarity & Relatedness presentations by Marten Postma, Minh Le, Alessandro Lopopolo, and Emiel van Miltenburg.
July 01, 2015 at VU Amsterdam
— VU University scientists cluster National Research Agenda
Led by Piek Vossen, a group of scientists at VU University automatically divided 11,700 questions from NWO’s National Research Agenda into clusters. On the basis of language technology and mathematical equations of the most important words, slightly over 60 clusters of questions were found which at their turn were classified in a few hundred sub-clusters. Important themes are health and energy, but also big data, art, and sports. NWO is happy with this analysis. The VU Topic Browser allows NWO to quickly and efficiently process the large number of responses.
— Mar. 13 2015: Why linguists are needed by George Lakoff — webcast available
Webcast available of presentation by George Lakoff on Friday March 13 2015 “Why linguists are needed: The severe limitations of big data analysis of linguistic corpora” arguing that “big data statistical methods by themselves were hopeless” in a multi-million dollar project on analyzing the conceptual metaphors in a vast corpus of US intelligence documents.
— NLP analysis of “The art of the Humanities – Graduation day 2014”
Introduction
— Open Source Dutch WordNet @CLIN 2015
At CLIN 2015, we presented Open Source Dutch Wordnet: (odwn_clin_2015)
The project website can be found at: project website.
— Similarity, co-occurrence, functional relation, part-whole relation, subcategorization, what else?
In word sense disambiguation and named-entity disambiguation, an important assumption is that a document consists of related concepts and entities.
— SensEval/Semeval output from participant systems (WSD)
If you are interested in having the individual output at token level for all the participant systems in the last SensEval/SemEval WSD tasks, we can find then now in a simple and homogeneous XML format, easy to process. You will find more information in our results section or in https://github.com/rubenIzquierdo/sval_systems
— DBPEDIA spotlight for KAF/NAF
If you are interested in extracting entities and their link to dbpedia entries, you should take a look to this module: https://github.com/rubenIzquierdo/dbpedia_ner It allows you to use a KAF or a NAF file with jus tokens and terms, calls to the DBPEDIA online webservice and extract entities and the link to dbpedia automatically.
— Sense annotated corpora in NAF
If you work in WSD or you are simply interested in sense annotated corpora, you should take a look at this GitHub repository. You will find some well-known corpora widely used within WSD task, manually annotated with WordNet senses and converted in our NAF format, which makes very easy the use of all our modules and pipelines. Currently these corpora are available:
— Demo Wsd4Kids
The Wsd4Kids demo implements a very simple Word Sense Disambiguation system and a graphical interface to interact with the system. The WSD system behind is based on a machine learning engine (Support Vector Machines) and uses a bag-of-word feature model.
— ULM1 participated at SemEval-2015 task #13: Multilingual WSD and Entity Linking
We started applying our ideas about the usage of the background information to develop a system able to perform WSD and Entity linking by using this kind of information. The task number 13 of the SemEval competition forum was selected to test our hypohesis: Multilingual WSD and Entity Linking
— ULM-3 participates in SemEval 2015 Task 4 TimeLine: Cross-Document Event Ordering
Members of the ULM-3 subproject participated in the SemEval 2015 Task 4 TimeLine: Cross-Document Event Ordering.
— Jun. 05 2015: ExProM Workshop
Workshop Extra-Propositional Aspects of Meaning (ExProM) in Computational Linguistics
— NLP for Policy Making
As a joint initiative with the Digital Humanites Group at Fondazione Bruno Kessler (FBK, Trento), we set up a collaboration with the Italian Ministry of Education, Universities and Research (MIUR) on the automatic analysis of linguistic data contained in the answers given by the participants of the public consultation “La Buona Scuola” (#labuonascuola).
— Jul. 2015: The First Workshop on Computing News Storylines [NewsStory]
In conjunction with ACL-IJCNLP 2015
July 2015, Bejing, China
— Oct. 30 2014: Round table on ‘Time and Language’
Event date:
Thursday, 30 October, 2014 – 18:30 to 20:00
— Oct. 23 2014: Reference Machine: visit Tech Labs Network Institute
Thursday October 23rd 2014: Piek, Alessandro, Emiel, Filip and Minh met up with Marco Otte and Desmond Germans in the Tech Labs of the Network Institute to exchange first ideas about building the Reference Machine.
Jun. 10, 2013: Piek Vossen winner of Spinoza Prize 2013
NWO Spinoza laureates announce plans for their prize money
// September 27, 2013
Can Machines Understand Language?
Understanding language by machines
1st VU-Spinoza workshop
October 17th 2014
— 1st VU-Spinoza Workshop ULM-1
ULM-1: The borders of ambiguity.
Presentation by: dr. Rubén Izquierdo Beviá & Marten Postma, MA.
— 1st VU-Spinoza Workshop ULM-2
ULM-2: Word, concept, perception and brain.
Presentation by: Emiel van Miltenburg, MA & Alessandro Lopopolo, MA.
— 1st VU-Spinoza Workshop ULM-3
ULM-3: Stories and world views as a key to understanding language.
Presentation by: dr. Tommaso Caselli & dr. Roser Morante Vallejo.
— 1st VU-Spinoza Workshop ULM-4
ULM-4: A quantum model of text understanding.
Presentation by: Minh Ngoc Lê, MSc & Filip Ilievski, Student assistant.
The borders of ambiguity
Within ULM-1, Rubén Izquierdo & Marten Postma focus on Word Sense Disambiguation (WSD).
Word, concept and the perception of images and sounds
Within ULM-2, Emiel van Miltenburg focusses on the perception of images and sounds.
Storylines and perspectives
Within ULM-3, Tommaso Caselli & Roser Morante & Chantal van Son work on stories and world views as a key to understanding language.
A quantum model of text understanding
Within ULM-4, Filip Ilievski and Minh Lê investigate a new model of natural-language-processing (NLP). ULM-4 consists of two sub projects carried out by two PhD candidates. Filip Ilievski works on defining context and background knowledge while Minh Lê investigates NLP architectures for background knowledge integration.


