Skip to main content

How links and websites Scans Work

Overview

Dralvia phishing scans are passive. The scanner collects signals about a URL or domain, assigns a score, and returns evidence for review.

1. Normalize the input

  • Normalize the URL or domain.
  • Validate the type (URL, domain, IP, email, or hash).

2. Collect signals

  • DNS, MX, and SPF data
  • WHOIS age and registrar metadata
  • Public web history (see below)
  • TLS certificate details and header posture
  • Certificate Transparency history (see below)
  • Host and infrastructure context (see below)
  • Redirect chains and landing behavior
  • Threat feeds (when configured)
  • Browser sandbox observations when a full review is needed

Findings that are recorded but do not count against a site

Not every observation in a report is evidence of phishing. Some are simply worth knowing, and a report that quietly hid them would be less useful, not more.

A few findings are therefore shown with a full explanation but carry no weight in the score and no effect on the recommended action. They matter only alongside real evidence. Today these are:

  • A hidden or zero-size iframe on the page. Consent banners, analytics, chat widgets and payment frames are all built this way. Measured across sites we track as known good and sites later confirmed malicious by independent threat feeds, this appears on far more legitimate pages than malicious ones.
  • A sign-in prompt. Asking you to log in is what most sites do.
  • Writing to browser storage. Cookies, localStorage and similar. Almost every modern site does this.

Each of these says so in its own explanation in the report, so you can tell at a glance which findings moved the score and which are context. If you are reviewing a site yourself, they are still worth reading: a hidden frame on a page that also asks for your password is a different thing from a hidden frame on a documentation site.

Public web history

WHOIS tells you when a domain was registered. Public web history tells you how long the domain has actually had a visible presence on the web. Dralvia reads the public Internet Archive history for the domain and records:

  • the earliest archived capture (first-seen),
  • how many distinct days of captures exist (footprint depth).

Why this matters: a domain can be registered for years yet have almost no public web history, and a brand-new lure domain has neither. A thin or absent public footprint is a corroborating signal of newly stood-up or recently repurposed infrastructure. It stacks with the registration-age signal rather than replacing it, and is surfaced as history:thin_public_footprint (with a recent first-seen date recorded as history:recent_first_seen for context). This lookup is best-effort: if the archive is slow or unreachable, the scan continues and simply omits the history.

Certificate Transparency history

Every public TLS certificate is logged to Certificate Transparency. Dralvia reads that public history for the domain and adds it to the evidence:

  • the number of certificates seen for the domain,
  • how recently the newest certificate was issued,
  • how long ago the earliest certificate appeared,
  • related names that share the same certificates,
  • whether a wildcard certificate is in use.

Why this matters: a brand-new certificate on its own is not suspicious, because legitimate sites renew certificates often. The signal Dralvia adds is the certificate footprint. When a domain's entire certificate history only just appeared, that points to newly stood-up infrastructure rather than a routine renewal, which is a common pattern for fresh phishing sites. Scans surface this as the ct:new_certificate_footprint finding, with the recent issuance recorded as ct:freshly_issued_certificate for context.

This lookup is best-effort. If the Certificate Transparency source is slow or unreachable, the scan continues normally and simply omits the certificate history, so it never delays a result.

Hosting provider and ASN

Dralvia identifies the hosting provider and ASN (the network that hosts the address a domain resolves to) and records it as named evidence on every scan, so a reviewer always sees who hosts the target. When the provider or ASN matches a marker that is disproportionately associated with phishing and malware staging (for example bulletproof or offshore hosting), the scan raises infrastructure:abuse_prone_hosting, which carries a modest weight and names the matched marker in the evidence. Simply identifying a mainstream provider is recorded as infrastructure:hosting_provider_identified and adds no score.

Host and infrastructure context

Dralvia records a fingerprint of the host that backs a domain so that two different lure pages running on the same infrastructure can be recognized later. Two parts:

  • Keyless by default: the TLS certificate fingerprint and the favicon hash are always recorded, with no extra setup.
  • Optional vendor context: when your workspace provides a Censys or Shodan key, Dralvia also records open ports, observed services, and the hosting or ASN organization for the host. Without a key, no vendor lookup is made.

The vendor lookup is best-effort: if the vendor is slow or unreachable, the scan continues and simply omits those fields.

Brand impersonation on the certificate

A benign-looking domain can ship a TLS certificate that also covers a brand-lookalike name. Dralvia compares the related names found on the certificate (via Certificate Transparency) against your trusted-brand catalogue, so a name that closely imitates a brand without being that brand raises brand:ct_sibling_impersonation_<brand>. This catches impersonation that the main domain on its own would hide.

The trusted-brand catalogue is loaded from configured brand records. If those records are empty, the scanner falls back to Dralvia's packaged brand catalogue so brand checks do not silently disappear.

Redirects are also checked after they resolve. If a short link or wrapper lands on a final host that structurally impersonates a tracked brand, such as a brand name placed under an unrelated registrable domain, the result carries the same brand impersonation evidence and can feed correlation scoring.

Addresses built to read as a different domain

Some addresses are constructed so the part you recognise is not the part that is registered. purolator.com-utoa.shop is registered as com-utoa.shop; the familiar purolator.com is just a label sitting in front of it. On a phone, where the address bar truncates, the familiar part may be all you see.

Dralvia points this out wherever it appears — in the address you scanned or in the address a redirect lands on — and names the domain that is actually registered. It does not depend on a brand catalogue, so it works for names nobody has listed.

This is context, not a finding, and it does not change your score. Legitimate organisations have addresses shaped the same way: net-entreprises.fr and info-retraite.fr are French public services, co-operative.coop is a century-old retailer. Measured against a million ranked domains, we could not find a reliable way to tell the two apart from the address alone — a new registration looks the same as an old one whose certificate was renewed yesterday, and most registries no longer publish registration dates at all.

So we tell you what the address is really registered as and leave the judgement with you, rather than scoring a guess. If that address arrived in a message you were not expecting, the mismatch is worth knowing about.

A short link or wrapper is judged on its destination, not on the wrapper. When a scan follows a redirect onto a different registered domain, brand findings raised about that final address are carried back onto the link you scanned.

Fixed on 11 August 2026

One shape was missing from that hand-off: an address that takes the brand's exact name and puts it under a different suffix — roblox.com.kz, paypal.com.pt. A link landing on the weaker roblox-login.com reported a brand finding, while a link landing on roblox.com.kz reported none, which is the wrong way round.

Both are now carried back. If a scan of a short link came back without a brand finding before this date, rescan it.

This adds context, not a conviction. An exact brand name under another suffix is also the shape of a brand's own country domain, so on its own it raises the finding and some score — it does not by itself produce an Avoid.

Redirect chains seen by the browser

A link can pass through several addresses before a page appears. Dralvia also uses the redirects observed by the browser, including those that happen before the destination loads, when assessing that chain. This can reveal a suspicious delivery route even when the first page contains little more than a redirect.

The findings explain whether the chain crosses several registered domains, carries many recognized affiliate or click-tracking parameters, or combines both patterns. A finding seen in more than one part of the scan counts once. Ordinary campaign parameters such as utm_source do not count as these tracker signals. The redirect summary retains recognized parameter names without their values.

These checks concern the page's navigation. Redirects for images, scripts, and embedded frames do not become part of its delivery route. Adult or dating content alone does not establish phishing or credential theft.

Review the redirect findings and Why this score in your scan result. Submit a new scan to collect current behavior; an older saved result keeps its original score. No additional setup is needed. The browser may observe only part of a chain within the scan's limits, and a site can send different visitors to different destinations, so the result describes what the scan observed.

The brand named has to be the right one

When a report says an address is imitating a brand, that name is the point. A report naming the wrong brand is worse than one naming none, so the rule for matching a brand inside a longer name is strict: the brand has to appear as a word of its own, not merely as letters buried inside another word.

Fixed on 12 August 2026

Until this date the match was a plain letter-search, so a short brand name could hide inside an ordinary English word. support, portal and suport all contain port, and the catalogue contains the University of Portsmouth (port.ac.uk). Five scam addresses were therefore reported as imitating a university, and one as imitating a college of medicine.

Measured against a million ranked addresses, requiring a whole word removes 194 such mistakes out of 199 and keeps every genuine lookalike.

Two consequences worth knowing:

  • Where an address really was imitating something, the correct brand is still named — and where two brands were listed, the wrong one is now dropped and the right one kept.
  • Two addresses moved from Avoid to Caution. Their only brand evidence had been the mistaken match, so once it went there was nothing else against them. A Caution still warns you. We would rather tell you less than tell you something untrue.
Added on 12 August 2026

A handful of well-known short forms are now matched: amz and amzn for Amazon, msft for Microsoft, goog for Google, nflx for Netflix. A short form only counts when it stands as a word of its own and the address also carries a lure word such as support, login or verify — support-v1amz.com is named as imitating Amazon, while amz.run and amzn.to are not touched.

Short forms are matched conservatively on purpose. Three letters coincide far more easily than a whole brand name, so this raises a finding worth showing but never enough on its own to reach Avoid, and it is never matched inside a longer word — amzsurprises does not count, for the same reason support no longer counts as port.

Other abbreviations, and short forms appearing only in the part of an address before the domain, are still not matched. Both need measurement we do not yet have, and a rule that guesses would flag ordinary sites.

Which brands these checks can protect

Every check in this section names a brand, and a brand can only be named if it is in the trusted-brand catalogue. A lookalike of a brand that is not in the catalogue raises no brand: finding, however obvious the imitation looks to a person.

That is worth stating plainly because the catalogue is not a list of "all important brands" — it grew from the customers and sectors Dralvia has served. On 11 August 2026 it held no crypto-wallet brand at all.

Uphold has been added. Measured against a million ranked domains it produces no false matches, and on addresses Dralvia had already scanned it now names seven that previously carried no brand finding.

Ledger, Exodus and Trezor were deliberately not added, and the reason is worth knowing if you were expecting them:

  • ledger and exodus are ordinary English words. Adding them would flag ledger-enquirer.com (a newspaper), sql-ledger.org, hledger.org and securityledger.com, among others.
  • trezor means "treasury" in Serbian and Czech, so adding it would flag trezor.gov.rs, a government address, and treezor.com, an unrelated French banking company.

A brand whose name is also an everyday word cannot be protected by name-shape matching without accusing legitimate sites. Those brands are still covered by every check that does not depend on the catalogue — certificate age, hosting, page behaviour, threat intelligence and campaign infrastructure.

If a brand matters to your organisation and is missing, it can be added to your own trusted-brand records.

A brand name misspelled behind another word

A misspelling is different from the shape above, because it is hard to explain innocently. Dralvia already flagged an address whose whole name is a brand with one character changed (whatapp.com for WhatsApp, paypa1.com for PayPal). It did not flag the same misspelling when a word was put in front of it, so login-microsft.com, secure-paypa1.com and audios-whatapp.cn passed with no brand finding at all. That is now brand:compound_typo_<brand>, and the report names the part of the address that is misspelled, not just the brand.

Two limits are deliberate:

  • It never decides a verdict by itself. It is worth fewer points than the same misspelling standing alone, because a word next to it might legitimately explain it — and it cannot reach Avoid without other evidence.
  • Ordinary words are not misspellings. Some everyday words sit one character from a brand name: finance from Binance, orange from OrangeX, krypto from Crypto.com. Dralvia treats a word as ordinary when it recurs across many unrelated names, which real businesses do and a typosquat does not. Measured against a million ranked domains, the check matches 18 — and several of those turned out to be genuine lookalikes that happen to be popular.

A brand name run together with a lure word, with no hyphen

Dralvia has long flagged addresses that pair a brand with a word like login, secure or verifypaypal-verify.com, secure-paypal.com. It turned out that deleting the hyphen was enough to avoid the check entirely: paypalverify.com and securepaypal.com produced no brand finding at all. Of eight such pairs tested, all eight hyphenated forms were flagged and none of the glued forms were.

The reason it worked that way is a real trade-off, not an oversight. Plenty of ordinary words contain a brand name — pineapple, snapple, applewood — and simply flagging any name that contains a brand would have flagged 22% of the eight hundred most-linked sites on the web.

So the check now asks the rest of the name to corroborate: once the brand is taken out, what is left must itself be a lure word. paypalverify leaves "verify" and is flagged; pineapple leaves "pine" and is not. The match has to be exact, which is what keeps purchase from being read as the bank Chase. Brand names shorter than five characters are excluded from this check entirely, because a short sequence inside a longer word is usually a coincidence.

The lure words are no longer only English. They now include their equivalents in Spanish, Portuguese, French, German and Italian — the campaign that exposed this was Spanish. As with the wording check on page text, short ambiguous words are left out: conta sits inside "contact" and "container", so it is not treated as a signal on its own.

Measured against the eight hundred most-linked legitimate sites, this change flags no additional site.

Updated 17 August 2026. The rest of the name may now also be a second brand name, not only a lure word. appleicloudlogin.com was missed by the original rule: taking out apple leaves "icloudlogin", which is a brand followed by a lure, and a check that only knew lure words could not read it. Names of this shape are now flagged.

A lure word is still required somewhere in the name. Two brand names run together with nothing else, such as appleicloud.com or applemusic.com, are not flagged: companies pair their own names constantly, and that pattern is usually a defensive registration rather than an attack.

A name like this raises one finding, not one per brand it mentions, and the finding names the more specific brand. The evidence is a single fact, that the address is assembled from a brand and a lure, and counting it twice would push a name-based observation on its own into the highest risk band. A suspicious name is evidence; it is not a verdict by itself.

Measured again against the same eight hundred sites: still no additional site flagged.

Code that was deliberately made unreadable

Most sites compress their JavaScript to make pages load faster. That is called minification, it is completely normal, and Dralvia does not hold it against a site.

Obfuscation is different. An obfuscator does not make code smaller — it makes code unreadable on purpose. The clearest sign is a lookup table: instead of the script containing ordinary words, every word it needs is hidden in a list of numbers and fetched by position, so nothing meaningful appears anywhere in the file. A tool that compresses your code has no reason to do that.

Dralvia reports this as content:machine_obfuscated_js, separately from the older and much weaker content:obfuscated_js, and the report tells you what was found rather than just that something was:

39,804-character script built by an obfuscator: constant string decoder, 722 hex literals (18.1 per 1k chars), 32 of 32 identifiers meaningless

When Dralvia recognises the specific injection, the report names it and shows a short campaign fingerprint, so two sites hit by the same attack are visibly the same attack. When it does not recognise it, the fingerprint is shown anyway marked not yet identified — being unnamed is not the same as being harmless, and it still lets you match one report against another.

This most often means a site has been hacked and had code injected into its pages — the site itself is genuine, its owner is usually unaware, and visitors are the ones at risk. Because of that, this finding is one of the few that is not waived for long-established, well-configured domains: being an old, real business is exactly the profile of a site that gets compromised.

Two deliberate limits:

  • It never decides a verdict by itself. It cannot reach Avoid without other evidence.
  • Ordinary compressed code is not flagged. Measured against the eight hundred most-linked sites on the web — 266 of which serve large amounts of minified JavaScript — this finding fired on none of them.

Escaped text is not hidden code

A related, much weaker finding is content:obfuscated_js, which notes packed or escaped JavaScript. It used to be raised on nearly one in three ordinary websites, because it counted any \u escape sequence — and that is simply how pages write non-English text and emoji. Decoding what those escapes actually said found Russian and Vietnamese menu labels, and, on most sites, the emoji-support script that every WordPress installation ships.

It now counts only escapes of ordinary ASCII characters, which is what hiding code looks like: there is no reason to write a as an escape sequence except to make it unreadable. Escaped accented text and emoji no longer count at all, and structured data (the block a site uses to describe its page to search engines) is no longer read as code.

That change removed this finding from 87% of the sites that previously carried it, with no loss on any confirmed compromised site — those are reported under the stronger obfuscation finding above.

A page copied off another site with a copying tool

Some scam sites are not built, they are copied. Offline website-copying tools leave their own marker in every page they write, and one of the common ones also records the address it copied from:

Mirrored from zeexpress-oficial.shop/ by HTTrack Website Copier ...

When Dralvia finds that marker it reports content:offline_mirror_artifact, and names the site that was copied.

Copying a site is not by itself wrong — mirroring your own site for an offline archive is a normal thing to do, and that case is scored lightly. What changes the picture is when the tool records that it copied a different domain from the one now serving the page: that address is serving somebody else's website under its own name, according to the tool that made it. That is content:cloned_from_other_host, and it is weighted accordingly.

This works for brands Dralvia has never heard of, because it does not consult a brand list at all — the page identifies itself.

Being the destination scam pages point at

campaign:known_redirect_target covers two different situations, and they deserve different treatment:

  • This address forwards visitors to a destination Dralvia has already flagged. That is evidence about the page you asked about, and it is always reported.
  • Several separate flagged pages were seen sending people here. For a well-known site, that is what being impersonated looks like, not what being malicious looks like.
Changed on 11 August 2026

Until this date the second situation was reported at full weight unless the site happened to be in Dralvia's brand catalogue. Because that catalogue did not list them, github.com was returned as Caution 53 and vercel.com as Caution 40 — almost entirely from this one finding — while google.com and facebook.com were not, purely because they were on the list.

A well-known destination is now recognised from Dralvia's own signals instead: its registration age, its certificate, and its own scan history. No list of names is involved, so the same protection reaches sites nobody has added.

This does not soften anything else. A site that forwards visitors into a scam campaign is still reported at full weight however established it is, and a page on a shared hosting platform cannot claim the platform's age to qualify. Of the 1,092 addresses carrying this finding in the last 60 days, 489 are bare IP addresses, which can never qualify.

Reused infrastructure

Dralvia remembers the certificate and favicon fingerprints of hosts it has flagged before. When a new scan shows a fingerprint that one or more previously-flagged hosts also used, the scan raises infrastructure:reused_known_bad_fingerprint, a strong signal that the host runs on reused attacker infrastructure even when the domain itself is brand new. The match is learned only from your own scan history, and a host is never penalized by its own past result.

A fingerprint has to be of something. Some servers answer a request for /favicon.ico with 200 OK and an empty body — this is common when the scan target is a bare IP address rather than a domain name. Every such host produces the same hash as every other one, because they are all the hash of nothing.

Dralvia does not treat that as shared infrastructure. An empty icon produces no fingerprint at all, and the scan keeps looking at the other icon paths, so a real icon served from /favicon.png is still found. The hash of an empty response is also rejected outright by the reputation lookup, so it can never link two hosts even if it reaches the store another way. The scan still records that the icon was empty, under favicon_empty in the technical evidence — the observation is kept, it just is not used as evidence of a relationship.

Two hosts that both serve nothing have not been shown to have anything in common.

Being on the same platform is not shared infrastructure. Hosting services like GitHub Pages, Vercel and Netlify serve every site they host from a single wildcard certificate — one certificate covering *.github.io, for example. That certificate is shared by every tenant of the service, most of whom have never heard of each other.

Dralvia does not treat a wildcard certificate belonging to a hosting platform as evidence of reused attacker infrastructure. Without this, the legitimate home page of a hosting service would be flagged because of the sites its own customers publish.

A certificate that is not a platform wildcard still counts, including a wildcard an attacker obtained for their own domain — that really is one operator's infrastructure being reused. And a site hosted on one of these platforms can still be flagged on any other evidence, including a shared site icon, so running a phishing kit on a large host does not hide it.

URL-level threat intelligence

Some malicious reports are tied to one exact path, not the whole site. For example, a host can serve a normal homepage while one deep link serves a risky download or a credential page. Dralvia checks threat intelligence against the full URL when the scan input includes a path or query string, then falls back to host-level matching when only a domain is available.

Your domain has to be the host, not just a word in the address. Threat feeds list full URLs, and a URL can mention a domain name somewhere other than its host — in a folder name, in a mirror path, or in a redirect parameter. A file at https://raw.githubusercontent.com/someone/example.com/setup.zip is hosted on raw.githubusercontent.com; example.com there is just part of a folder name somebody chose.

Dralvia reads the host out of each listed URL and only counts entries actually hosted on the domain being scanned. A domain is never flagged because its name appears inside somebody else's link. The listed URLs shown as evidence are filtered the same way, so what you see is only what your own domain is listed for.

When an exact URL is listed in a configured threat source, the result can include a threatfeed:* finding. The explanation names the finding family and the score shows it as threat intelligence evidence. Stale or unreachable feed data does not create a finding by itself. If feed coverage is unavailable, the scan continues and records the degraded coverage in the technical evidence.

When a platform hosts someone else's malicious file

The major URL threat feeds list full URLs, not domains. Anyone can upload a file to a code host, a file locker or a chat CDN, and if that file is malicious the feed lists its URL — which sits under the platform's domain.

Asking "is any URL on this host listed?" is the right question for a throwaway phishing domain and the wrong one for a platform where the public uploads files. Left uncorrected it convicts the platform itself: on 5 August 2026 that was github.com (1,084 listed URLs), raw.githubusercontent.com (5,754), drive.google.com (246), and five other well-known services.

When you scan the site itself, Dralvia now separates the two cases. The full threat-intelligence weight is kept unless all of the following hold:

  • the listing matched the host, not the exact URL you asked about;
  • every listed entry that belongs to that host points at a path — a file somebody published — rather than naming the host itself;
  • you asked about the site, with no path or query of your own;
  • the domain is established on Dralvia's own signals: its age and certificate, or a clean history in your own past scans. Never by name, and never from a list of approved brands. A clean history means scans that actually reached the site, spread over at least five separate days and a month or more; see Domain Reputation.

In that case the finding becomes threatfeed:hosted_malicious_url, a smaller signal that says plainly: malicious files have been hosted on this domain, and the domain itself remains usable. It is never removed, because it is true and you should see it.

The full weight still applies when you scan the listed URL itself, when you scan any page with a path, when the feed names the host rather than files under it, when a second feed lists the host without that evidence, and for any domain that is not established — a domain registered days ago and already serving malware is exactly the threat this is for.

This applies to every domain, including the widely recognised ones. A recognisable name is not a pass: if a threat feed names such a host directly, rather than files published under it, the scan still returns Avoid at full weight. A compromised well-known service is the single most useful thing a scan can tell you, and nothing in this rule softens it.

Two limits worth knowing:

  • A subdomain does not inherit its parent's standing. raw.githubusercontent.com and drive.google.com are scored on their own evidence, and neither has a registration age of its own, so both still return Avoid on the feed listing. This is deliberate: anyone can open a tenant on a shared platform in a minute, and letting a subdomain borrow the platform's age is how a lure gets treated as established.
  • The rule needs a domain we can age or have scanned before. A platform we have no history for keeps the full weight.
Corrected 6 August 2026

This rule shipped on 5 August but did not take effect; the platforms it describes were scored by a path that never consulted it. Between those two dates a scan of one of them could still return Avoid. Rescan to get the current result — a scan you ran in that window may still be showing you its cached verdict.

Provider verdicts such as Google Safe Browsing are unaffected: those judge a domain, not a URL.

Browser sandbox behavior

For full reviews, Dralvia can render the page in an isolated browser sandbox and turn high-confidence observations into scoring evidence. Examples include:

  • a risky download that the sandbox blocks,
  • a login form paired with brand text on a host that does not belong to that brand,
  • a hosting provider warning page that says the URL was reported for phishing,
  • a page on free or cloud hosting that shows a login surface and makes external browser requests,
  • an off-brand page that calls a known brand's browser API endpoints.

These observations appear as content:* findings in the result. They are not based on the page merely using a common provider or lacking a security header. The finding is added only when the rendered behavior supplies stronger evidence.

A script file that the page loads as part of working normally is not treated as a download. Only a file the browser is actually navigating to, or a program-style file (.exe, .msi, .apk, and similar), counts as a blocked download.

Pages hosted on someone else's platform

Anyone can open a free account on a helpdesk, notes or site-builder platform and publish a page under that platform's domain — for example some-tenant.freshdesk.com or some-kit.github.io. The page inherits the platform's domain age, its certificate and its respectable-looking address, none of which the author had to earn.

Dralvia treats those pages as what they are:

  • Age is not inherited. A page on a shared platform does not get the "long established domain" benefit from the platform's registration date. It can still earn a good standing from its own scan history on that exact address.
Corrected 11 August 2026

This was written as a general rule, but it was enforced from a shorter list of platforms than the one the scanner uses to recognise shared hosting in the first place. On twenty-one platforms — including weebly.com, dropboxusercontent.com, ipfs.io, 000webhostapp.com, framer.website, dweb.link, godaddysites.com, neocities.org and the InfinityFree family (epizy.com, rf.gd, 42web.io, lovestoblog.com, great-site.net, iceiy.com, wuaze.com, infinityfreeapp.com) — a page did inherit the platform's registration date, which could soften its verdict.

The two lists are now one list, so any platform the scan recognises as shared hosting is also one whose age a tenant cannot borrow. If you scanned a page on one of those platforms before this date, rescan it: the earlier verdict may have been more lenient than it should have been.

  • A call to action that leaves the platform is a finding. If such a page's main button or link ("View document", "Sign in", "Verify") sends you to a completely unrelated domain, the scan raises content:offsite_credential_cta. This is a common phishing split: the trustworthy-looking page carries the story, and the page that asks for your details is somewhere else entirely.

Neither applies to an ordinary website on its own domain. A company linking from its own site to its own document portal is normal and is not flagged.

Lure wording, and the languages it is checked in

Some pages ask you to "verify your account", "confirm your identity" or pass a "security check". Dralvia looks for that wording — but only reports it when the page also has somewhere to put your details: a password field, a form that submits to another site, or a button that hands you off to one.

That condition matters. Plenty of ordinary sites use those exact words in help articles and security pages; a news site explaining a security check is not asking you for anything. Without the condition, checking the top 800 most-linked sites on the web flagged eight of them — news sites, a jobs board, a hardware maker — purely for writing about account security. With it, one remains, and that one really does post its form to another host.

The same wording is now checked in Portuguese, Spanish, French, German and Italian as well as English. Scam pages are written for the people they target, and just over a fifth of the pages Dralvia examines that state a language are not in English. Only complete phrases are matched, never single words: conta, cuenta, compte and konto are as ordinary in those languages as "account" is in English.

Brand names found in page text

When a page already looks suspicious, Dralvia also notes which well-known brand names appear in its visible text, as content:brand_logo_<brand>. It is a supporting observation, not a verdict: a page that asks for a password while talking about a bank it has nothing to do with is worth a second look.

This is a text match, not an image match. Despite the flag's name, nothing here inspects a logo. Dralvia is looking for the brand's name as a word in the page's visible wording. Matching a brand's actual logo image is a separate check and reports separately. The flag name is a historical one that is being corrected.

Only the words a visitor can read count. The page's underlying code is not searched, because brand names appear there constantly for reasons that have nothing to do with the page's content: an analytics tag names its vendor, an icon file is called apple-touch-icon.png, a sharing widget carries the name of every network it supports, and every page on the web opens with a meta element. None of that is visible to anyone reading the page, so none of it can be an attempt to pass the page off as somebody else's.

Specifically excluded: element attributes, HTML comments, scripts and stylesheets. If the page is too broken to parse, the code is stripped of its tags and the remaining words are searched instead, so a malformed page is still checked rather than skipped.

Each brand is reported once. A large company often has several domains on Dralvia's brand list, such as a company's main site alongside its country sites. Naming that company once on a page is one finding, reported under the company's primary domain, not one finding per domain on the list. Dralvia reports at most three brands per page, and that limit now counts brands rather than list entries, so one company can no longer take every slot and hide the others.

Both boundaries were corrected on 2026-08-22. Before that the check searched the page's code as well as its words, which produced brand names for pages that never mentioned the brand anywhere a reader could see, and it counted a company once per list entry. Scores on affected pages go down, never up, and no other check changed. If you have a saved report from before that date, re-run the scan to see the current result.

Three things it deliberately does not do:

  • A site is never flagged for naming itself, including its own other domains. A company's own domain showing its own brand is not impersonation, and neither is one of its sibling properties: google.com mentioning "gmail", or a brand's country domain mentioning the brand. Dralvia decides this from the brand registry, which records which domains belong to the same company — so it applies only to domains that are listed there. A lookalike such as brand.com.xx is not listed, cannot inherit an owner, and is still flagged, whatever it looks like.

  • Ordinary words are not treated as brands. The brand list also serves as a reference set of known-good sites, and includes many universities whose short names are everyday English words. Matching those found prose rather than impersonation, so they are excluded from the text search. That exclusion covers academic registries worldwide — .edu, and .ac.uk, .edu.au, .ac.jp, .edu.gh and their equivalents — not only US universities. Impersonation of those organisations is still detected by the lookalike-domain checks, which compare the address rather than the wording.

  • A long-established site is not accused of impersonation for naming another brand. Most sign-in pages on the web offer "Sign in with Google" or "Continue with Apple", and most of them belong to companies that are not on any 380-entry brand list. Treating that as impersonation punished sites for a gap in Dralvia's own reference data rather than for anything they did. So when a domain is established — either Dralvia's own scan history shows it repeatedly clean, or it has been registered for over two years and presents a valid certificate — a brand name appearing in its text no longer counts as evidence of impersonation.

    This is narrow on purpose. Being established does not excuse a misspelled or lookalike address, a homograph, or serving another brand's actual favicon: those all still carry full weight at any age. And a domain hosted as a tenant on a shared platform cannot borrow that platform's registration age. Newly registered credential-collection pages are unaffected — of the pages this rule changed, every single one that lost an impersonation finding was a long-established, genuine business, and no phishing detection was lost.

Sites that refuse to be analyzed

Many legitimate sites challenge automated visitors, and a challenge from a recognized security provider (Cloudflare, Akamai, Imperva and similar) is treated as the mild, common signal it is: content:scanner_blocked.

A different case is a site that refuses the scan and, in place of the page, returns a short custom holding message such as "Preparing secure session" while running code that stores identifiers in your browser. That is not a standard bot challenge — it is a bespoke screen deciding who is allowed to see the real content, which is how a phishing page stays hidden from analysis while still opening for the person who was sent the link. Dralvia raises content:analysis_evasion_gate for it, and it is enough on its own to move a result to Caution.

The finding requires all of: the site blocked the scan, no recognized provider challenge was present, the response contained no real page content, and the holding page was doing something to the visitor (browser storage, a redirect, or a permission prompt). A site that returns a real page, even behind a block, does not raise it.

3. Score and explain

  • Combine weighted signals into a risk score.
  • Attach explanations and flags for every failed check.

4. Produce evidence

  • Return the full result payload.
  • Generate an EvidencePack when enabled.
  • Write a transparency log entry when enabled.

Limits and ethics

Only scan assets you own or are authorized to assess. Dralvia does not exploit targets or perform intrusive testing.