科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Data in brief2026-08-01

PhishVN: A time-stamped Vietnamese URL phishing dataset with impersonation-scenario labels and confidence tiers.

Thai Nguyen Vu

原始摘要(英文原文)· Original abstract
PhishVN is an open phishing-website dataset localised to the Vietnamese context, compiled from public national and community threat feeds and allow-lists. Its core is a time-stamped, verified URL table of 18,997 records (2587 phishing from the national feed, 16,410 legitimate), extended by an explicitly-tagged bronze expansion stratum, community-reported domains from the ChongLuaDao project (with per-domain first-seen dates reconstructed for ∼ 50 % from public archival traces) and recent OpenPhish URLs, for 53,116 records in total, each described by a 21-feature lexical/infrastructure schema that is aligned with the CompPhish dataset to enable cross-dataset study. Phishing URLs are drawn from the National Cyber Security Centre feed (Tin Nhiem Mang); legitimate URLs form three explicitly-tagged strata of increasing difficulty: the certified "trusted organisation" directory (easy, curated), a .vn-filtered Tranco slice (popular sites Vietnamese users actually visit), and a global Tranco sample (hard, traffic-based negatives that break the "non-.vn implies phishing" shortcut), to strengthen external validity. Every record carries an impersonation-scenario label (bank, government, tax, e-commerce, telecom, delivery, social, gaming) inferred from brand tokens in the URL, and a label-confidence tier(gold/silver/bronze) derived from the source's verification status. Records carrying a source-attested event date (the national-feed core phishing, the certified organisations and the date-reconstructed half of the bronze stratum) are ordered temporally and split with grouping by registrable domain, so that near-duplicate campaign subdomains never span train and test; records whose registrable-domain group carries no attested event date (most of the web-whitelist and top-list negatives, and a small undated remainder of the bronze stratum) are placed into splits by grouped random assignment. We release the datasheet, the full collection and processing pipeline, and reproducible URL baseline detectors reported with multi-seed spreads and bootstrap confidence intervals. A preliminary multi-modal companion additionally pairs 868 URLs across both classes (209 phishing, 659 benign) with the rendered page HTML and a screenshot; the text-message and email channels are left to a planned extension.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

PhishVN: A time-stamped Vietnamese URL phishing dataset with impersonation-scenario labels and confidence tiers. — 科研速览 Science Skim