The Mannheim Search Join Engine


Lehmberg, Oliver ; Ritze, Dominique ; Ristoski, Petar ; Meusel, Robert ; Paulheim, Heiko ; Bizer, Christian



DOI: https://doi.org/10.1016/j.websem.2015.05.001
URL: http://www.sciencedirect.com/science/article/pii/S...
Weitere URL: http://dl.acm.org/citation.cfm?id=2870130
Dokumenttyp: Zeitschriftenartikel
Erscheinungsjahr: 2015
Titel einer Zeitschrift oder einer Reihe: Web Semantics
Band/Volume: 35
Heft/Issue: 3
Seitenbereich: 159-166
Ort der Veröffentlichung: Amsterdam [u.a.]
Verlag: Elsevier
ISSN: 1570-8268
Sprache der Veröffentlichung: Englisch
Einrichtung: Fakultät für Wirtschaftsinformatik und Wirtschaftsmathematik > Information Systems V: Web-based Systems (Bizer 2012-)
Fakultät für Wirtschaftsinformatik und Wirtschaftsmathematik > Web Data Mining (Juniorprofessur) (Paulheim 2013-2017)
Fachgebiet: 004 Informatik
Freie Schlagwörter (Englisch): Table extension ; Data search ; Search Joins ; Web tables ; Microdata ; Linked data
Abstract: A Search Join is a join operation which extends a user-provided table with additional attributes based on a large corpus of heterogeneous data originating from the Web or corporate intranets. Search Joins are useful within a wide range of application scenarios: Imagine you are an analyst having a local table describing companies and you want to extend this table with attributes containing the headquarters, turnover, and revenue of each company. Or imagine you are a film enthusiast and want to extend a table describing films with attributes like director, genre, and release date of each film. This article presents the Mannheim Search Join Engine which automatically performs such table extension operations based on a large corpus of Web data. Given a local table, the Mannheim Search Join Engine searches the corpus for additional data describing the entities contained in the input table. The discovered data are joined with the local table and are consolidated using schema matching and data fusion techniques. As a result, the user is presented with an extended table and given the opportunity to examine the provenance of the added data. We evaluate the Mannheim Search Join Engine using heterogeneous data originating from over one million different websites. The data corpus consists of HTML tables, as well as Linked Data and Microdata annotations which are converted into tabular form. Our experiments show that the Mannheim Search Join Engine achieves a coverage close to 100% and a precision of around 90% for the tasks of extending tables describing cities, companies, countries, drugs, books, films, and songs.




Dieser Eintrag ist Teil der Universitätsbibliographie.




Metadaten-Export


Zitation


+ Suche Autoren in

+ Aufruf-Statistik

Aufrufe im letzten Jahr

Detaillierte Angaben



Sie haben einen Fehler gefunden? Teilen Sie uns Ihren Korrekturwunsch bitte hier mit: E-Mail


Actions (login required)

Eintrag anzeigen Eintrag anzeigen