Focused Crawling for Structured Data
Meusel, Robert
;
Mika, Peter
;
Blanco, Roi

DOI:
|
https://doi.org/10.1145/2661829.2661902
|
URL:
|
https://s.yimg.com/ge/labs/v2/uploads/anthelion.pd...
|
Additional URL:
|
http://de.slideshare.net/RobertMeusel/focused-craw...
|
Document Type:
|
Conference or workshop publication
|
Year of publication:
|
2014
|
Book title:
|
CIKM 2014 : Proceedings of the 23rd ACM International Conference on Conference on Information and Knowledge Management
|
Page range:
|
1039-1048
|
Location of the conference venue:
|
Shanghai, China
|
Date of the conference:
|
November 3-7, 2014
|
Place of publication:
|
New York, NY
|
Publishing house:
|
ACM
|
ISBN:
|
978-1-4503-2598-1
|
Publication language:
|
English
|
Institution:
|
School of Business Informatics and Mathematics > Wirtschaftsinformatik V (Bizer)
|
Subject:
|
004 Computer science, internet
|
Keywords (English):
|
bandit-based selection , focused crawling , microdata , online learning
|
Abstract:
|
The Web is rapidly transforming from a pure document collection to the largest connected public data space. Semantic annotations of web pages make it notably easier to extract and reuse data and are increasingly used by both search engines and social media sites to provide better search experiences through rich snippets, faceted search, task completion, etc. In our work, we study the novel problem of crawling structured data embedded inside HTML pages. We describe Anthelion, the first focused crawler addressing this task. We propose new methods of focused crawling specifically designed for collecting data-rich pages with greater efficiency. In particular, we propose a novel combination of online learning and bandit-based explore/exploit approaches to predict data-rich web pages based on the context of the page as well as using feedback from the extraction of metadata from previously seen pages. We show that these techniques significantly outperform state-of-the-art approaches for focused crawling, measured as the ratio of relevant pages and non-relevant pages collected within a given budget.
|
 | Dieser Eintrag ist Teil der Universitätsbibliographie. |
Search Authors in
You have found an error? Please let us know about your desired correction here: E-Mail
Actions (login required)
 |
Show item |
|
|