Tapioca: Tapioca is a search engine for topically similar RDF datasets.

Tapioca is a search engine for finding topically similar linked data datasets.

Demo Issues Source Code

The Web of data is growing continuously with respect to both the size and number of the datasets published. Porting these datasets to five-star Linked Data however requires data publishers to link their novel dataset with the already available Linked Data sets. Given the size and growth of the Linked Data Cloud, the current mostly manual approach used for detecting relevant datasets for linking is thus obsolete.

We present Tapioca, a linked dataset search engine so as to provide data publishers with similar existing datasets automatically. Our search engine uses a novel approach for determining the topical similarity of datasets. This approach relies on probabilistic topic modelling to determine related datasets by relying solely on the metadata of datasets.

The source code can be found at Github. The software is provided under a dual license. For non-commercial purposes, the terms of the LGPL 3.0 license hold. For commercial purposes, please contact us.

For our publication Detecting Similar Linked Datasets Using Topic Modelling we have the following additional material:

  • For the first experiment, you can find the gold standard as well as the detailed F1 scores of Tapioca and a second version of Tapioca that uses the Jensen-Shannon divergence, in this folder.
  • For the second experiment, you can find the detailed values of the P(w|T) and the A measure in this folder.
  • For the third experiment, you can find the detailed values of the P(w|T) and the A measure as well as the F1 scores of our approach in this folder.

Project Team

Publications

by (Editors: ) [BibTex of ]

News

AKSW at web.br in São Paulo ( 2018-10-22T09:37:49+02:00 by Natanael Arndt)

2018-10-22T09:37:49+02:00 by Natanael Arndt

From October 1st until 6th a delegation from AKSW Group, Leipzig University of Applied Sciences (HTWK), eccenca GmbH, and Max Planck Institute for Human Cognitive and Brain Sciences went to São Paulo, Brazil to meet people from the Web Technologies … Continue reading → Read more about "AKSW at web.br in São Paulo"

AskNow 0.1 Released ( 2018-09-13T15:35:04+02:00 by Prof. Dr. Jens Lehmann)

2018-09-13T15:35:04+02:00 by Prof. Dr. Jens Lehmann

Dear all, we are very happy to announce AskNow 0.1 – the initial release of Question Answering Components and Tools over RDF Knowledge Graphs. Website: http://asknow.sda.tech/ Demo: http://asknowdemo.sda.tech GitHub: https://github. Read more about "AskNow 0.1 Released"

Jekyll RDF Tutorial Screencast ( 2018-08-07T11:11:12+02:00 by Natanael Arndt)

2018-08-07T11:11:12+02:00 by Natanael Arndt

Since 2016 we are developing Jekyll-RDF a plugin for the famous Jekyll–static website generator. Read more about "Jekyll RDF Tutorial Screencast"

DBpedia Day @ SEMANTiCS 2018 ( 2018-07-20T14:37:25+02:00 by Johannes Frey)

2018-07-20T14:37:25+02:00 by Johannes Frey

Don’t miss the 12th edition of the DBpedia Community Meeting in Vienna, the city with the highest quality of life in the world. Read more about "DBpedia Day @ SEMANTiCS 2018"

SANSA 0.4 (Semantic Analytics Stack) Released ( 2018-06-26T18:33:38+02:00 by Prof. Dr. Jens Lehmann)

2018-06-26T18:33:38+02:00 by Prof. Dr. Jens Lehmann

We are happy to announce SANSA 0.4 – the fourth release of the Scalable Semantic Analytics Stack. Read more about "SANSA 0.4 (Semantic Analytics Stack) Released"