Content
Phishing attackers need victims to click. That means the URL has to look convincing enough to pass a quick glance. But creating a convincing URL is harder than it seems.
Legitimate domains are short and clean: paypal.com, chase.com, github.com. Phishing domains, by contrast, must work around the fact that they don't own the real domain. They register something close, bury the real brand name in a subdomain, or add keywords that create false context.
Security researchers have identified a short list of URL rules that catch a surprising number of phishing attempts: suspicious words like "login," "verify," "secure," "OTP," and "update"; odd brand spellings like paypa1 instead of paypal; long or deep subdomains; excessive hyphens; high digit or symbol ratios; and cases where a brand string appears only in a subdomain and not in the registered domain .
These aren't random patterns. They reflect the constraints attackers face.
Building the Core Detection Logic
Let's walk through the key checks your tool should perform. Each one targets a specific phishing tactic.
Check 1: Suspicious Keywords in the URL
Phishing pages often include urgency and action words in the URL to make the link feel transactional and legitimate. Common examples include login, verify, secure, account, update, confirm, and password.
text
SUSPICIOUS_WORDS = ['login', 'verify', 'secure', 'account',
'update', 'confirm', 'password', 'otp']
def has_suspicious_words(url):
url_lower = url.lower()
return [word for word in SUSPICIOUS_WORDS if word in url_lower]
These words appear in legitimate URLs too—plenty of real sites have /login or /account paths. That's why this is a flag, not a verdict. Your tool should note the presence and let the human interpret context.
Check 2: Misleading Subdomains
This is one of the most effective phishing tricks. Attackers register a domain they control and then place a trusted brand name in a subdomain:
text
https://paypal.com.secure-login.evil-site.net/account
The real domain here is evil-site.net. Everything before it is a subdomain. But a quick glance might register "paypal.com" and assume safety.
Your code should parse the URL and extract the registered domain—the actual domain the site owner controls. Everything to the left is just subdomain decoration.
python
from urllib.parse import urlparse
def get_domain_parts(url):
parsed = urlparse(url)
hostname = parsed.hostname or ''
parts = hostname.split('.')
return {
'full_host': hostname,
'subdomain': '.'.join(parts[:-2]) if len(parts) > 2 else '',
'registered_domain': '.'.join(parts[-2:]) if len(parts) >= 2 else hostname
}
If a known brand appears in the subdomain but not in the registered domain, that's a strong signal of deception .
Check 3: Brand Imitation and Typosquatting
Attackers register domains that look like legitimate brands with small changes:
paypa1.com (digit "1" replacing "l")
arnazon.com (rn resembling m)
faceb00k.com (zeros replacing o's)
google-verify.com (hyphenated brand + keyword)
A simple approach is comparing the registered domain against a list of common brands using edit distance. If the domain is close but not exact, flag it.
python
def levenshtein_distance(s1, s2):
if len(s1) < len(s2):
return levenshtein_distance(s2, s1)
if len(s2) == 0:
return len(s1)
previous_row = range(len(s2) + 1)
for i, c1 in enumerate(s1):
current_row = [i + 1]
for j, c2 in enumerate(s2):
insertions = previous_row[j + 1] + 1
deletions = current_row[j] + 1
substitutions = previous_row[j] + (c1 != c2)
current_row.append(min(insertions, deletions, substitutions))
previous_row = current_row
return previous_row[-1]
This won't catch every spoof—determined attackers use Unicode characters that look identical to Latin letters—but it catches the common cases.
Check 4: URL Structure Anomalies
Phishing URLs tend to be longer, deeper, and messier than legitimate ones. Features that matter include:
Total URL length (phishing URLs are often 75+ characters)
Number of dots (deep subdomains)
Number of hyphens (brand-keyword combinations)
Presence of an IP address instead of a domain name
Excessive path depth (/a/b/c/d/login.html)
python
def structure_anomalies(url):
parsed = urlparse(url)
hostname = parsed.hostname or ''
flags = []
if len(url) > 75:
flags.append('long_url')
if hostname.count('.') > 3:
flags.append('deep_subdomains')
if url.count('-') > 2:
flags.append('many_hyphens')
if parsed.path.count('/') > 3:
flags.append('deep_path')
return flags
A tool called LinkSentry, which uses 111 URL features, found that directory length, URL length, and slash count were among the top five most important predictors of phishing . Structure matters.
Turning Flags into a Score (Not a Verdict)
A good detection tool doesn't say "this is phishing." It says "this URL has 4 suspicious indicators." The difference matters because every single indicator produces false positives.
Legitimate sites use "secure" and "login" in URLs. Real companies have long subdomains. Hyphens appear in plenty of safe domains.
The goal is risk scoring, not binary judgment. You might weight indicators differently:
Brand appearing only in subdomain: high weight
IP address instead of domain: high weight
Suspicious keyword + brand imitation together: elevated
Long URL alone: low weight
When multiple indicators stack up, confidence increases. A single flag means "look closer." Three or four flags means "don't enter credentials here."
The Limits of URL Analysis
Your tool won't catch everything. Sophisticated attackers use clean, short domains that look legitimate. They can register domains that contain no suspicious keywords and no brand imitation. They use HTTPS with valid certificates because certificates are free and easy to obtain .
URL analysis is one layer. It's fast, it's explainable, and it catches the majority of low-effort phishing attacks. But it should be combined with other signals: domain age (new domains are riskier), page content analysis, and visual comparison against known legitimate sites .
The most important thing your tool does is slow you down. Phishing succeeds when you click without thinking. A tool that flags suspicious patterns interrupts that automatic behavior—and that interruption is often enough to prevent the mistake.