Predictive Recall
Dralvia's scanner reads public threat feeds while it scans. That makes feed membership a poor way to grade it: if a site is already listed, the scanner was handed the answer, so agreeing with the feed proves nothing.
Predictive recall removes that problem by grading only the cases where the answer did not exist yet.
The rule
A scan counts toward predictive recall only when it ran before any feed had ever listed that domain. In those cases the listing cannot have influenced the verdict, because it did not exist. What Dralvia said was a genuine prediction, and the listing that arrived later grades it.
Everything else is discarded. A scan that ran after a listing is thrown out, not scored, because it cannot tell you anything you did not already know.
Lead time changes the meaning
The gap between the scan and the first listing matters, so results are reported in three groups rather than as one number.
| Gap between scan and first listing | What a "Safe" verdict there means |
|---|---|
| Up to 7 days | Most likely a genuine miss. This is the headline figure. |
| 8 to 30 days | Mixed. Some misses, some sites that changed. |
| More than 30 days | Usually a site that was genuinely fine when scanned and was compromised or repurposed later. |
Collapsing these into one number would be misleading. A site that was clean in May and turned malicious in August is not a detection failure, and counting it as one would understate the scanner while explaining nothing about why.
Recall is never published alone
Any detector can reach perfect recall by warning about everything. So every measurement also scores a fixed list of well known, legitimate sites, and reports how many of those were warned about.
The two numbers are always published together. A change that lifts recall by warning about ordinary websites shows up immediately in the second number, and both are alerted on.
The published figure covers recent listings only
A verdict is graded against the scan that produced it, and that scan is history. Once a domain has been listed, nothing Dralvia does later can change how it was judged at the time. A lifetime average would therefore be held down permanently by older versions of the scanner and would barely move when detection improves.
The headline figure covers domains a feed first listed in the last 30 days. Wider windows and the all-time figure are published alongside it, so the direction of travel is visible and nothing is hidden by the narrower view.
Honest gaps in the method
The known-good list ages. Those sites are graded on their most recent stored scan, which can be weeks or months old. A verdict may already have been corrected without the measurement noticing. Every result therefore reports the age of the oldest scan behind it, so the figure is never read as more current than it is.
The sample is not a random draw. Domains enter this measurement because Dralvia scanned them before they were listed, and Dralvia scans candidates that look worth checking. One day's batch can dominate a window and score better than the rest of it. Every result therefore publishes its own composition: how many domains, across how many distinct days, and what share came from the largest single day. A figure drawn mostly from one batch is reported as such.
Feeds contain more than phishing. Public feeds also list developer demo projects, test pages, and short lived clones. Domains identified this way are reviewed before any of them is treated as a defect.
No measurement is not the same as a bad measurement. When there is too little qualifying data, no recall figure is published at all. A rate is never reported as zero to fill a gap, because "we caught none of them" and "we have not measured yet" are opposite statements.
What this does not cover
Predictive recall measures false negatives: sites Dralvia should have warned about and did not. It says nothing on its own about how quickly a verdict is produced or how a site is categorised. Those are covered in the detection quality methodology.