스톰크롤러
StormCrawler| 개발자 | DigitalPebble, Ltd. |
|---|---|
| 초기 릴리즈 | 2014년 9월 11일 |
| 안정적 해제 | 2.2 / 2022년 1월 11일; 전 |
| 리포지토리 | |
| 기록 위치 | 자바 |
| 유형 | 웹 크롤러 |
| 면허증 | 아파치 라이선스 |
| 웹사이트 | 스톰크롤러 |
StormCrawler는 Apache Storm에 대기 시간이 짧고 확장 가능한 웹 크롤러를 구축하기 위한 오픈 소스 리소스 모음입니다.아파치 라이선스로 제공되며 주로 자바(프로그래밍 언어)로 작성된다.
StormCrawler는 모듈형으로 핵심 모듈로 구성되어 있으며, 가져오기, 구문 분석, URL 필터링과 같은 웹 크롤러의 기본 구성 요소를 제공한다.핵심 부품과는 별도로, 이 프로젝트는 Elasticsearch 및 Apache Solr를 위한 스푸트 및 볼트나 Apache Tika를 사용하여 다양한 문서 형식을 구문 분석하는 파서 볼트 같은 외부 자원도 제공한다.
이 프로젝트는 여러 회사의 생산에 사용된다.[1]
아마존닷컴은 2016년 10월 스톰크롤러의 저자와 함께 Q&A를 출간했다.[2]인포큐는 2016년 12월에 1대를 운영했다.[3]아파치너치와의 비교 벤치마크는 2017년 1월 dzone.com에 게재됐다.[4]
몇몇 연구 논문에서는 특히 다음과 같이 StormCrawler의 사용을 언급했다.
위키프로젝트에는 온라인에서 이용 가능한 동영상과 슬라이드 목록이 포함되어 있다.[8]
StormCrawler는 특히 Common Crawler에[9] 의해 공개적으로 이용 가능한 대규모 뉴스 데이터 세트를 생성하는 데 사용된다.
참고 항목
참조
- ^ "Powered By · DigitalPebble/storm-crawler Wiki · GitHub". Github.com. 2017-03-02. Retrieved 2017-04-19.
- ^ "StormCrawler: An Open Source SDK for Building Web Crawlers with ApacheStorm Linux.com The source for Linux information". Linux.com. 2016-10-12. Retrieved 2017-04-19.
- ^ "Julien Nioche on StormCrawler, Open-Source Crawler Pipelines Backed by Apache Storm". Infoq.com. 2016-12-15. Retrieved 2017-04-19.
- ^ "The Battle of the Crawlers: Apache Nutch vs. StormCrawler - DZone Big Data". Dzone.com. Retrieved 2017-04-19.
- ^ "Crawling the German Health Web: Exploratory Study and Graph Analysis".
- ^ "MirasText: An Automatically Generated Text Corpus for Persian".
- ^ Sanagavarapu, Lalit Mohan; Mathur, Neeraj; Agrawal, Shriyansh; Reddy, Y. Raghu (2018). Advances in Information Retrieval. Lecture Notes in Computer Science. Vol. 10772. pp. 811–814. doi:10.1007/978-3-319-76941-7_81. ISBN 978-3-319-76940-0.
- ^ "Presentations · DigitalPebble/storm-crawler Wiki · GitHub". Github.com. 2017-04-04. Retrieved 2017-04-19.
- ^ "News Dataset Available – Common Crawl".