A Novel Web Scraping Approach Using the Additional Information Obtained from Web Pages

Uzun, Erdinç

dc.contributor.author	Uzun, Erdinç
dc.date.accessioned	2022-05-11T14:02:59Z
dc.date.available	2022-05-11T14:02:59Z
dc.date.issued	2020
dc.identifier.issn	2169-3536
dc.identifier.uri	https://doi.org/10.1109/ACCESS.2020.2984503
dc.identifier.uri	https://hdl.handle.net/20.500.11776/4566
dc.description.abstract	Web scraping is a process of extracting valuable and interesting text information from web pages. Most of the current studies targeting this task are mostly about automated web data extraction. In the extraction process, these studies first create a DOM tree and then access the necessary data through this tree. The construction process of this tree increases the time cost depending on the data structure of the DOM Tree. In the current web scraping literature, it is observed that time efficiency is ignored. This study proposes a novel approach, namely UzunExt, which extracts content quickly using the string methods and additional information without creating a DOM Tree. The string methods consist of the following consecutive steps: searching for a given pattern, then calculating the number of closing HTML elements for this pattern, and finally extracting content for the pattern. In the crawling process, our approach collects the additional information, including the starting position for enhancing the searching process, the number of inner tag for improving the extraction process, and tag repetition for terminating the extraction process. The string methods of this novel approach are about 60 times faster than extracting with the DOM-based method. Moreover, using these additional information improves extraction time by 2.35 times compared to using only the string methods. Furthermore, this approach can easily be adapted to other DOM-based studies/parsers in this task to enhance their time efficiencies. © 2013 IEEE.	en_US
dc.language.iso	eng	en_US
dc.publisher	Institute of Electrical and Electronics Engineers Inc.	en_US
dc.identifier.doi	10.1109/ACCESS.2020.2984503
dc.rights	info:eu-repo/semantics/openAccess	en_US
dc.subject	algorithm design and analysis	en_US
dc.subject	Computational efficiency	en_US
dc.subject	document object model	en_US
dc.subject	web crawling and scraping	en_US
dc.subject	Efficiency	en_US
dc.subject	Extraction	en_US
dc.subject	Information use	en_US
dc.subject	Trees (mathematics)	en_US
dc.subject	Websites	en_US
dc.subject	Construction process	en_US
dc.subject	Extraction process	en_US
dc.subject	Extraction time	en_US
dc.subject	String methods	en_US
dc.subject	Text information	en_US
dc.subject	Time efficiencies	en_US
dc.subject	Web data extraction	en_US
dc.subject	Web scrapings	en_US
dc.subject	Data mining	en_US
dc.title	A Novel Web Scraping Approach Using the Additional Information Obtained from Web Pages	en_US
dc.type	article	en_US
dc.relation.ispartof	IEEE Access	en_US
dc.department	Fakülteler, Çorlu Mühendislik Fakültesi, Bilgisayar Mühendisliği Bölümü	en_US
dc.identifier.volume	8	en_US
dc.identifier.startpage	61726	en_US
dc.identifier.endpage	61740	en_US
dc.institutionauthor	Uzun, Erdinç
dc.relation.publicationcategory	Makale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı	en_US
dc.authorscopusid	54783608800
dc.identifier.wos	WOS:000527413400002	en_US
dc.identifier.scopus	2-s2.0-85083419796	en_US

Bu öğenin dosyaları:

Ad:: 4566.pdf
Boyut:: 4.949Mb
Biçim:: PDF
Açıklama:: Tam Metin / Full Text

Göster/Aç

Bu öğe aşağıdaki koleksiyon(lar)da görünmektedir.

Scopus İndeksli Yayınlar Koleksiyonu [4328]
Scopus Indexed Publications Collection
WoS İndeksli Yayınlar Koleksiyonu [4789]
WoS Indexed Publications Collection
Çorlu Mühendislik Fakültesi Koleksiyonu [990]

Basit öğe kaydını göster