Larascraper

Larascraper is a scraping toolkit for Laravel. You write a class, describe what you want, and it runs through a real headless browser or plain HTTP without changing the shape of your code.
Written against version 3.3.
Installation
composer require edulazaro/larascraper
The browser driver uses Puppeteer, so the first time you scrape a JavaScript-heavy site it pulls the browser down. The HTTP driver needs nothing extra.
Your first scraper
A scraper is a class with a handle() method. Inside it, $this->scrape($url) starts a fetch and a terminal decides what comes back.
use EduLazaro\Larascraper\Scraper; use EduLazaro\Larascraper\ScraperResponse; class BikeScraper extends Scraper { protected function handle(string $url): ScraperResponse { return $this->scrape($url)->run(); } }
Run it from anywhere:
$response = BikeScraper::run('https://example.com/bikes/4'); $response->data; // the raw HTML $response->success; // true or false $response->error; // the error message if it failed
Parsing with a Crawler
A Crawler turns a page into data, through a small query builder over the DOM:
use EduLazaro\Larascraper\Crawler; class BikeCrawler extends Crawler { protected function handle(): array { return [ 'name' => $this->filter('h1.title')->text(), 'price' => $this->filter('.price')->text(), 'specs' => $this->filter('.spec')->each(fn ($node) => $node->text()), ]; } }
Chain it onto the fetch:
return $this->scrape($url)->crawl(BikeCrawler::class)->run(); // ['name' => 'Black Eagle', 'price' => '2,499', 'specs' => [...]]
The Crawler input is generic. Usually it is an HTML string, but $this->filter($selector, 'xml') parses as XML, and $this->raw() hands you the untouched input for regex, simplexml or json_decode, handy when an endpoint returns XML or JSON dressed up as a page.
Browser actions
Larascraper runs a real browser by default, and you can act on the page before grabbing the HTML.
return $this->scrape($url) ->type('#search', 'black eagle') ->press('Enter', waitForNavigation: true) ->waitForSelector('.results') ->crawl(BikeCrawler::class) ->run();
Actions run inside the browser, in order, and then the crawler sees the final DOM. The full set: click(), select() (single value or array for multi-selects), check() and uncheck(), hover(), scroll(), waitForSelector(), waitForNavigation().
For a page that only needs plain HTTP:
$this->scrape($url)->driver('http')->run();
Same code shape, no browser. Actions only make sense with the browser, so combining them with the HTTP driver throws.
Optional waits
A wait that times out normally fails the scrape. When an empty result set is a legitimate outcome rather than an error:
->waitForSelector('.results', ['optional' => true, 'timeout' => 8000])
The run continues and your crawler returns zero rows. You can also wait for whichever of several things lands first:
->waitForSelector(['.results', '.no-results'])
Retrying until a condition holds
use EduLazaro\Larascraper\Support\Condition; $this->scrape($url) ->repeatUntil( Condition::selectorMissing('#captcha'), fn ($b) => $b ->solveCaptcha('#captcha-img', '#captcha-input') ->clickAndWait('#verify'), max: 5, delay: 1500, ) ->crawl(ResultCrawler::class) ->run();
Always bounded, so it can never hammer a server. A throw inside one attempt counts as a failed attempt rather than a dead run: the loop re-checks the condition and tries again, up to max. A real configuration mistake (a bad captcha solver name, a 4xx from an API key) aborts immediately instead of retrying pointlessly.
Captchas
Simple image captchas go through OCR, with no API cost:
->solveCaptcha('#captcha-img', '#captcha-input')
The OCR packages are optional:
php artisan larascraper:install --captcha
For distorted ones that tesseract chokes on there is an OpenAI vision solver, opted into per call, reading OPENAI_API_KEY:
->solveCaptcha('#captcha-img', '#captcha-input', ['solver' => 'vision'])
OCR stays the default, so you only pay for the hard ones.
Downloading files
$response = $this->scrape($url) ->clickAndWait('a.download') ->capture(['expect' => 'application/pdf']) ->file(); $response->file; // the bytes $response->contentType; // application/pdf
Spiders
One scraper fetches one page. For thousands, a Spider calls your scrapers and pool() runs them concurrently.
use EduLazaro\Larascraper\Spider; class CatalogSpider extends Spider { protected int $concurrency = 20; public function handle(): array { $ids = range(1, 5000); return $this->pool( $ids, BikeScraper::class, fn ($response, $id) => $response->success ? $response->data : null, $this->concurrency, ); } } $bikes = array_filter(CatalogSpider::run());
PHP is single-threaded, so this uses Fibers: each scraper suspends when it fetches, the Spider gathers the in-flight requests into one wave and sends them together, then resumes each Fiber with its response.
Real concurrency lives on the HTTP driver. A shared cookie Session threads through every request, so a login survives the whole crawl.
built and maintained by Edu Lazaro · MIT license