Pythonic Crawling / Scraping Framework based on Non Blocking I/O operations.

Last update: Dec 05, 2022

Related tags

Web Crawling crawley

Overview

Pythonic Crawling / Scraping Framework Built on Eventlet

Features

High Speed WebCrawler built on Eventlet.
Supports relational databases engines like Postgre, Mysql, Oracle, Sqlite.
Supports NoSQL databased like Mongodb and Couchdb. New!
Export your data into Json, XML or CSV formats. New!
Command line tools.
Extract data using your favourite tool. XPath or Pyquery (A Jquery-like library for python).
Cookie Handlers.
Very easy to use (see the example).

Documentation

http://packages.python.org/crawley/

Project WebSite

http://project.crawley-cloud.com/

To install crawley run

~$ python setup.py install

or from pip

~$ pip install crawley

To start a new project run

~$ crawley startproject [project_name]
~$ cd [project_name]

Write your Models

""" models.py """

from crawley.persistance import Entity, UrlEntity, Field, Unicode

class Package(Entity):
    
    #add your table fields here
    updated = Field(Unicode(255))    
    package = Field(Unicode(255))
    description = Field(Unicode(255))

Write your Scrapers

""" crawlers.py """

from crawley.crawlers import BaseCrawler
from crawley.scrapers import BaseScraper
from crawley.extractors import XPathExtractor
from models import *

class pypiScraper(BaseScraper):
    
    #specify the urls that can be scraped by this class
    matching_urls = ["%"]
    
    def scrape(self, response):
                        
        #getting the current document's url.
        current_url = response.url        
        #getting the html table.
        table = response.html.xpath("/html/body/div[5]/div/div/div[3]/table")[0]
        
        #for rows 1 to n-1
        for tr in table[1:-1]:
                        
            #obtaining the searched html inside the rows
            td_updated = tr[0]
            td_package = tr[1]
            package_link = td_package[0]
            td_description = tr[2]
            
            #storing data in Packages table
            Package(updated=td_updated.text, package=package_link.text, description=td_description.text)


class pypiCrawler(BaseCrawler):
    
    #add your starting urls here
    start_urls = ["http://pypi.python.org/pypi"]
    
    #add your scraper classes here    
    scrapers = [pypiScraper]
    
    #specify you maximum crawling depth level    
    max_depth = 0
    
    #select your favourite HTML parsing tool
    extractor = XPathExtractor

Configure your settings

""" settings.py """

import os 
PATH = os.path.dirname(os.path.abspath(__file__))

#Don't change this if you don't have renamed the project
PROJECT_NAME = "pypi"
PROJECT_ROOT = os.path.join(PATH, PROJECT_NAME)

DATABASE_ENGINE = 'sqlite'     
DATABASE_NAME = 'pypi'  
DATABASE_USER = ''             
DATABASE_PASSWORD = ''         
DATABASE_HOST = ''             
DATABASE_PORT = ''     

SHOW_DEBUG_INFO = True

Finally, just run the crawler

~$ crawley run

Pythonic Crawling / Scraping Framework based on Non Blocking I/O operations.

Related tags

Overview

Pythonic Crawling / Scraping Framework Built on Eventlet

Features

Documentation

Project WebSite

To install crawley run

or from pip

To start a new project run

Write your Models

Write your Scrapers

Configure your settings

Finally, just run the crawler

Owner

Juan Manuel Garcia

A Telegram crawler to search groups and channels automatically and collect any type of data from them.

Nekopoi scraper using python3

Demonstration on how to use async python to control multiple playwright browsers for web-scraping

Quick Project made to help scrape Lexile and Atos(AR) levels from ISBN

抢京东茅台脚本，定时自动触发，自动预约，自动停止

This Scrapy project uses Redis and Kafka to create a distributed on demand scraping cluster

SmartScraper: 简单、自动、快捷的Python网络爬虫

A Python module to bypass Cloudflare's anti-bot page.

This program will help you to properly scrape all data from a specific website

A multithreaded tool for searching and downloading images from popular search engines. It is straightforward to set up and run!

A Python Covid-19 cases tracker that scrapes data off the web and presents the number of Cases, Recovered Cases, and Deaths that occurred because of the pandemic.

a way to scrape a database of all of the isef projects

京东茅台抢购最新优化版本，京东秒杀，添加误差时间调整，优化了茅台抢购进程队列

TarkovScrappy - A nifty little bot that lets you know if a queried item might be required for a quest at some point in the land of Tarkov!

Consulta de CPF e CNPJ na Receita Federal com Web-Scraping

A Python Oriented tool to Scrap WhatsApp Group Link using Google Dork it Scraps Whatsapp Group Links From Google Results And Gives Working Links.

This is a web crawler that works on employ email data by gmane.org and visualizes it in different ways.

A repository with scraping code and soccer dataset from understat.com.

✂️🕷️ Spider-Cut is a Network Mapper Framework (NMAP Framework)

This program scrapes information and images for movies and TV shows.