DEER: RDF Data Extraction and Enrichment Framework

Over the last years, the Linked Data principles have been used across academia and industry to publish and consume structured data. Thanks to the fourth Linked Data principle, many of the RDF datasets used within these applications contain implicit and explicit references to more data. For example, music datasets such as Jamendo include references to locations of record labels, places where artists were born or have been, etc. Datasets such as Drugbank contain references to drugs from DBpedia, were verbal description of the drugs and their usage is explicitly available. The goal of mapping component, dubbed DEER, is to retrieve this information, make it explicit and integrate it into data sources according to the specifications of the user. To this end, DEER relies on a simple yet powerful pipeline system that consists of two main components: enrichment functions and operators.

Download Issues

Enrichment functions and operators.

Enrichment functions implement functionality for processing the content of a dataset (e.g., applying named entity recognition to a particular property). Thus, they take a dataset as input and return a dataset as output. Enrichment operators work at a higher level of granularity and combine datasets. Thus, they take sets of datasets as input and return sets of datasets.

RDF specification paradigm

In the current version of DEER we introduce our new RDF based specification paradigm. The main idea behind this new paradigm is to enable the processing execution of specifications in an efficient way. To this end, we first decided to use RDF as language for the specification. This has the main advantage of allowing for creating specification repositories which can be queried easily with the aim of retrieving accurate specifications for the use cases at hand. Moreover, extensions of the specification language do not require a change of the specification language due to the intrinsic extensibility of ontologies. The third reason for choosing RDF as language for specifications is that we can easily check the specification for correctness by using a reasoner, as the specification ontology allows for specifying the restrictions that specifications must abide by.

Publications

by (Editors: ) [BibTex of ]

News

SANSA 0.2 (Semantic Analytics Stack) Released ( 2017-06-13T18:18:28+02:00 by Prof. Dr. Jens Lehmann)

2017-06-13T18:18:28+02:00 by Prof. Dr. Jens Lehmann

The AKSW and Smart Data Analytics groups are happy to announce SANSA 0.2 – the second release of the Scalable Semantic Analytics Stack. Read more about "SANSA 0.2 (Semantic Analytics Stack) Released"

AKSW at ESWC 2017 ( 2017-06-12T10:53:35+02:00 Christopher Schulz)

2017-06-12T10:53:35+02:00 Christopher Schulz

Hello Community! The ESWC 2017 just ended and we give a short report of the course at the conference, especially regarding the AKSW-Group. Our members Dr. Muhammad Saleem, Dr. Mohamed Ahmed Sherif, Claus Stadler, Michael Röder, Prof. Dr. Read more about "AKSW at ESWC 2017"

Four papers accepted at WI 2017 ( 2017-06-10T15:01:31+02:00 Christopher Schulz)

2017-06-10T15:01:31+02:00 Christopher Schulz

Hello Community! We proudly announce that The International Conference on Web Intelligence (WI) accepted four papers by our group. The WI takes place in Leipzig between the 23th – 26th of August. Read more about "Four papers accepted at WI 2017"

AKSW Colloquium, 29.05.2017, Addressing open Machine Translation problems with Linked Data. ( 2017-05-26T13:51:11+02:00 by Diego Moussallem)

2017-05-26T13:51:11+02:00 by Diego Moussallem

At the AKSW Colloquium, on Monday 29th of May 2017, 3 PM, Diego Moussallem will present two papers related to his topic. First paper titled “Using BabelNet to Improve OOV Coverage in SMT” of Du et al. Read more about "AKSW Colloquium, 29.05.2017, Addressing open Machine Translation problems with Linked Data."

SML-Bench 0.2 Released ( 2017-05-11T13:01:45+02:00 by Patrick Westphal)

2017-05-11T13:01:45+02:00 by Patrick Westphal

Dear all, we are happy to announce the 0.2 release of SML-Bench, our Structured Machine Learning benchmark framework. SML-Bench provides full benchmarking scenarios for inductive supervised machine learning covering different knowledge representation languages like OWL and Prolog. Read more about "SML-Bench 0.2 Released"