Gelişmiş Arama

Basit öğe kaydını göster

dc.contributor.authorUzun, Erdinç
dc.date.accessioned2022-05-11T14:02:59Z
dc.date.available2022-05-11T14:02:59Z
dc.date.issued2020
dc.identifier.issn2169-3536
dc.identifier.urihttps://doi.org/10.1109/ACCESS.2020.2984503
dc.identifier.urihttps://hdl.handle.net/20.500.11776/4566
dc.description.abstractWeb scraping is a process of extracting valuable and interesting text information from web pages. Most of the current studies targeting this task are mostly about automated web data extraction. In the extraction process, these studies first create a DOM tree and then access the necessary data through this tree. The construction process of this tree increases the time cost depending on the data structure of the DOM Tree. In the current web scraping literature, it is observed that time efficiency is ignored. This study proposes a novel approach, namely UzunExt, which extracts content quickly using the string methods and additional information without creating a DOM Tree. The string methods consist of the following consecutive steps: searching for a given pattern, then calculating the number of closing HTML elements for this pattern, and finally extracting content for the pattern. In the crawling process, our approach collects the additional information, including the starting position for enhancing the searching process, the number of inner tag for improving the extraction process, and tag repetition for terminating the extraction process. The string methods of this novel approach are about 60 times faster than extracting with the DOM-based method. Moreover, using these additional information improves extraction time by 2.35 times compared to using only the string methods. Furthermore, this approach can easily be adapted to other DOM-based studies/parsers in this task to enhance their time efficiencies. © 2013 IEEE.en_US
dc.language.isoengen_US
dc.publisherInstitute of Electrical and Electronics Engineers Inc.en_US
dc.identifier.doi10.1109/ACCESS.2020.2984503
dc.rightsinfo:eu-repo/semantics/openAccessen_US
dc.subjectalgorithm design and analysisen_US
dc.subjectComputational efficiencyen_US
dc.subjectdocument object modelen_US
dc.subjectweb crawling and scrapingen_US
dc.subjectEfficiencyen_US
dc.subjectExtractionen_US
dc.subjectInformation useen_US
dc.subjectTrees (mathematics)en_US
dc.subjectWebsitesen_US
dc.subjectConstruction processen_US
dc.subjectExtraction processen_US
dc.subjectExtraction timeen_US
dc.subjectString methodsen_US
dc.subjectText informationen_US
dc.subjectTime efficienciesen_US
dc.subjectWeb data extractionen_US
dc.subjectWeb scrapingsen_US
dc.subjectData miningen_US
dc.titleA Novel Web Scraping Approach Using the Additional Information Obtained from Web Pagesen_US
dc.typearticleen_US
dc.relation.ispartofIEEE Accessen_US
dc.departmentFakülteler, Çorlu Mühendislik Fakültesi, Bilgisayar Mühendisliği Bölümüen_US
dc.identifier.volume8en_US
dc.identifier.startpage61726en_US
dc.identifier.endpage61740en_US
dc.institutionauthorUzun, Erdinç
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanıen_US
dc.authorscopusid54783608800
dc.identifier.wosWOS:000527413400002en_US
dc.identifier.scopus2-s2.0-85083419796en_US


Bu öğenin dosyaları:

Thumbnail

Bu öğe aşağıdaki koleksiyon(lar)da görünmektedir.

Basit öğe kaydını göster