Laradomains

Laradomains answers the questions an application asks about a domain somebody typed: is it a real domain, which part of it was registered, how old is it, who is the registrar, does it take mail, is it known for malware. Every answer comes from a third party (the registry, a DNS resolver, the Wayback Machine) and never from the domain itself. That is the point: a domain pasted into a form may be hostile, and fetching it is the one thing you should not do.
It came out of Valor Web, which values websites without ever loading them, and ToxicFilter, which needs the age of the domain behind an email address and a clean way to tell a lookalike from an accented name.
Installation
composer require edulazaro/laradomains
PHP 8.4+ with the intl extension, Laravel 12 or newer. The config is optional: php artisan vendor:publish --tag=laradomains-config.
Five pieces
| Piece | Answers | Source |
|---|---|---|
Domain |
What is the host, the suffix, the registrable part, is it a lookalike | The Public Suffix List, bundled |
Domains::rdap() |
Registration and expiry dates, registrar, status, nameservers, DNSSEC | The registry's own RDAP server |
Domains::age() |
How old the domain is | RDAP, or the first Wayback capture |
Domains::dns() |
Addresses, MX, TXT, NS, whether it takes mail | DNS over HTTPS |
Domains::screen() |
Malware, phishing or adult content | Cloudflare's filtering resolvers |
Parsing
Domain::parse() takes what people paste (https://News.BBC.co.uk:8080/path, user@example.com, ñandú.es) and keeps the host, lower-cased, as ASCII for machines and Unicode for people. IP addresses are refused.
use EduLazaro\Laradomains\Domain; $domain = Domain::parse('https://News.BBC.co.uk:8080/path'); $domain->registrable(); // "bbc.co.uk" $domain->suffix(); // "co.uk" $domain->subdomain(); // "news" Domain::parse('ñandú.es')->ascii; // "xn--and-6ma2c.es"
The registrable part follows the Public Suffix List rather than "the last two labels", which is wrong for co.uk, gob.es or com.mx. Suffixes run by hosting platforms are opt-in: for the registry, edulazaro.github.io belongs to github.io; with registrable(private: true) it is a site of its own.
isLookalike() catches the two tricks phishing relies on: a label that mixes scripts (аpple.com with a Cyrillic "а") and a label written wholly in Cyrillic or Greek whose every letter has a Latin twin (аррӏе.com, all Cyrillic, which reads "apple"). An accented name in one script, ñandú.es, a real Cyrillic word, яндекс.com, and Cyrillic under .ru or .рф, where it is the norm, stay clean.
Domain::parse('аррӏе.com')->skeleton(); // "apple.com" Domain::parse('www.paypa1.com')->imitates(['paypal.com']); // "paypal.com"
skeleton() gives the Latin reading of a name and imitates() checks it against the brands you care about, 0/o and 1/l swaps included. impersonates() adds the shapes phishing actually uses, the brand as a label or a hyphenated part of someone else's domain: paypal.com.secure-login.io, paypal-secure.com. A brand that is a common word matches ordinary sites too (apple-pie-recipes.com), so it is a reason to review a link, not to block it. typosquats() catches the typing slips: paypall.com, payal.com, paypla.com, goog1e.com, arnazon.com with rn for m; but not paypay.com, a real company one ordinary letter away. typosquat() also says which slip it was, because they do not weigh the same: a lookalike letter is rarely an accident, a letter added or dropped is where ordinary words such as apples land, and kinds: keeps only the strong ones. And imitates() uses Unicode's full confusables data too, so gօօgle.com with an Armenian օ is recognised as Google.
Hosts are read the way a browser reads them, because that decides where a link really goes: a backslash ends the host, so https://evil.example\@paypal.com/login is evil.example.
RDAP without a middleman
RDAP replaced WHOIS, and it has no central database: every registry runs its own server, and IANA publishes which server answers for each TLD. Laradomains ships that file and asks the registry directly, so a lookup is one request, it needs no cache to be fast, and no relay sees the names your users type. A relay (rdap.org) is asked only for TLDs the IANA file does not list, and can be switched off.
use EduLazaro\Laradomains\Facades\Domains; $registration = Domains::rdap('news.bbc.co.uk'); // Nominet, about bbc.co.uk $registration->registeredAt; // 1994-12-13 $registration->registrar; $registration->expiresAt; $registration->status; // ["client transfer prohibited", ...]
Three outcomes are kept apart because each one means something different: supported false when the TLD has no RDAP at all (.es, .de, .io), registered false when the registry says nobody holds the name, and failed() when the lookup did not complete. Only the last one is worth retrying.
Age
Domains::age('example.com')->years(); // from the registry Domains::age('ejemplo.es', wayback: true)->source; // "wayback" where the registry has no date Domains::age('example.com')->isNewerThan(30); // registered this month?
Registries such as .es, .de or .io publish no RDAP, so there is no registration date to give. For those, certificates: true looks up the first certificate in the Certificate Transparency logs: a phishing domain usually gets one the day it is registered, which is exactly what a "registered less than a week ago" rule needs to see. The answer is read as a stream up to 2 MB, since any certificate read proves the domain existed by then, so even a domain with thousands of certificates answers within a short timeout.
The Wayback fallback is off by default. Its CDX server allows about a dozen requests a minute, which suits a queued job and not a sign-up form, and a first capture is a lower bound, not a registration date. With LARADOMAINS_AGE_CACHE_FOR set, ages are remembered in a persistent cache store, keyed by the registrable domain, and a failed lookup is kept for five minutes: a spam campaign repeating the same new domain costs one registry request, not one per message.
Keeping the lists current
Both lists the package reads, the Public Suffix List and the IANA RDAP bootstrap, ship with it. php artisan laradomains:update downloads fresh copies, checks they look like the real thing before replacing anything, and compiles the suffix list to a PHP array that opcache keeps in memory. Schedule it monthly.
DNS and screening
DNS goes over HTTPS by default, so the answers do not depend on whatever resolver the server happens to use:
Domains::dns()->acceptsMail('example.com'); // false with no MX or a null MX (RFC 7505) Domains::dns()->txt('example.com'); // multi-string records joined
Screening asks Cloudflare's filtering resolvers, which answer 0.0.0.0 for the names they block:
Domains::screen('example.com', timeout: 1.5); // "clean", "malware", "adult" or "unknown"
Both resolvers are asked in parallel, so screening takes as long as the slower one, typically well under 200 ms. It fails closed: when a resolver does not answer, the result is "unknown", never "clean", and the caller decides what that means. Domains::verdict() keeps the two answers apart, so a timeout on the adult resolver does not hide that the malware one already cleared the domain, and verdicts can be cached per host. The adult category is broad, catching piracy and cannabis shops too, so it is a flag for review, not proof.
Every service has its own timeout and retries (screening two seconds and none, Wayback twenty and a five-second pause), and every network call can pass its own timeout:, so a request path is never held by the slowest service.
Rate limits and tests
Every request first passes through a hook with the service name (rdap, dns, screen, wayback), the place for a rate limiter shared across workers:
Domains::beforeRequest(fn (string $service) => $throttle->wait($service));
And because every call goes through Laravel's HTTP client, Http::fake() covers the whole package in tests.
built and maintained by Edu Lazaro · MIT license