Focused Crawling for Structured Data


Meusel, Robert ; Mika, Peter ; Blanco, Roi



DOI: https://doi.org/10.1145/2661829.2661902
URL: https://s.yimg.com/ge/labs/v2/uploads/anthelion.pd...
Additional URL: http://de.slideshare.net/RobertMeusel/focused-craw...
Document Type: Conference or workshop publication
Year of publication: 2014
Book title: CIKM 2014 : Proceedings of the 23rd ACM International Conference on Conference on Information and Knowledge Management
Page range: 1039-1048
Location of the conference venue: Shanghai, China
Date of the conference: November 3-7, 2014
Place of publication: New York, NY
Publishing house: ACM
ISBN: 978-1-4503-2598-1
Publication language: English
Institution: School of Business Informatics and Mathematics > Information Systems V: Web-based Systems (Bizer 2012-)
Subject: 004 Computer science, internet
Keywords (English): bandit-based selection , focused crawling , microdata , online learning
Abstract: The Web is rapidly transforming from a pure document collection to the largest connected public data space. Semantic annotations of web pages make it notably easier to extract and reuse data and are increasingly used by both search engines and social media sites to provide better search experiences through rich snippets, faceted search, task completion, etc. In our work, we study the novel problem of crawling structured data embedded inside HTML pages. We describe Anthelion, the first focused crawler addressing this task. We propose new methods of focused crawling specifically designed for collecting data-rich pages with greater efficiency. In particular, we propose a novel combination of online learning and bandit-based explore/exploit approaches to predict data-rich web pages based on the context of the page as well as using feedback from the extraction of metadata from previously seen pages. We show that these techniques significantly outperform state-of-the-art approaches for focused crawling, measured as the ratio of relevant pages and non-relevant pages collected within a given budget.




Dieser Eintrag ist Teil der Universitätsbibliographie.




Metadata export


Citation


+ Search Authors in

+ Page Views

Hits per month over past year

Detailed information



You have found an error? Please let us know about your desired correction here: E-Mail


Actions (login required)

Show item Show item