Technology

The Future of Technology Starts Here

AI · Web3 · Cloud · Cyber · next-gen dev

Programming

How to Spot a Phishing Website Using Code

Description

Phishing websites are designed to look like legitimate login portals, but their URLs often betray them. This guide walks you through building a basic URL-analysis tool that flags suspicious domains, misleading subdomains, and unusual URL structures. You'll learn why these indicators are clues—not proof—of malicious intent, and how to use code to make smarter security decisions.

Introduction

You check the address bar, see the padlock, and assume you're safe. That assumption is exactly what phishing attackers count on.

HTTPS no longer signals trustworthiness. Over 62% of unique phishing URLs detected in a recent study used HTTPS protocol, and the number keeps growing . The padlock only means your connection is encrypted—it says nothing about who is on the other end .

So how do you spot a phishing site before you hand over credentials? One approach is to stop relying on visual trust cues and start reading the URL itself. Phishing URLs follow patterns: they mimic brand names, overload subdomains, use suspicious keywords, and hide the real domain in the middle of the string.

In this article, you'll learn how to build a basic URL-analysis tool that catches these patterns. More importantly, you'll understand why each indicator is a clue rather than proof—and how to interpret what your code finds.

Content

Why URLs Are a Goldmine for Phishing Detection
Phishing attackers need victims to click. That means the URL has to look convincing enough to pass a quick glance. But creating a convincing URL is harder than it seems.

Legitimate domains are short and clean: paypal.com, chase.com, github.com. Phishing domains, by contrast, must work around the fact that they don't own the real domain. They register something close, bury the real brand name in a subdomain, or add keywords that create false context.

Security researchers have identified a short list of URL rules that catch a surprising number of phishing attempts: suspicious words like "login," "verify," "secure," "OTP," and "update"; odd brand spellings like paypa1 instead of paypal; long or deep subdomains; excessive hyphens; high digit or symbol ratios; and cases where a brand string appears only in a subdomain and not in the registered domain .

These aren't random patterns. They reflect the constraints attackers face.

Building the Core Detection Logic
Let's walk through the key checks your tool should perform. Each one targets a specific phishing tactic.

Check 1: Suspicious Keywords in the URL

Phishing pages often include urgency and action words in the URL to make the link feel transactional and legitimate. Common examples include login, verify, secure, account, update, confirm, and password.

text
SUSPICIOUS_WORDS = ['login', 'verify', 'secure', 'account',
'update', 'confirm', 'password', 'otp']

def has_suspicious_words(url):
url_lower = url.lower()
return [word for word in SUSPICIOUS_WORDS if word in url_lower]
These words appear in legitimate URLs too—plenty of real sites have /login or /account paths. That's why this is a flag, not a verdict. Your tool should note the presence and let the human interpret context.

Check 2: Misleading Subdomains

This is one of the most effective phishing tricks. Attackers register a domain they control and then place a trusted brand name in a subdomain:

text
https://paypal.com.secure-login.evil-site.net/account
The real domain here is evil-site.net. Everything before it is a subdomain. But a quick glance might register "paypal.com" and assume safety.

Your code should parse the URL and extract the registered domain—the actual domain the site owner controls. Everything to the left is just subdomain decoration.

python
from urllib.parse import urlparse

def get_domain_parts(url):
parsed = urlparse(url)
hostname = parsed.hostname or ''
parts = hostname.split('.')
return {
'full_host': hostname,
'subdomain': '.'.join(parts[:-2]) if len(parts) > 2 else '',
'registered_domain': '.'.join(parts[-2:]) if len(parts) >= 2 else hostname
}
If a known brand appears in the subdomain but not in the registered domain, that's a strong signal of deception .

Check 3: Brand Imitation and Typosquatting

Attackers register domains that look like legitimate brands with small changes:

paypa1.com (digit "1" replacing "l")

arnazon.com (rn resembling m)

faceb00k.com (zeros replacing o's)

google-verify.com (hyphenated brand + keyword)

A simple approach is comparing the registered domain against a list of common brands using edit distance. If the domain is close but not exact, flag it.

python
def levenshtein_distance(s1, s2):
if len(s1) < len(s2):
return levenshtein_distance(s2, s1)
if len(s2) == 0:
return len(s1)
previous_row = range(len(s2) + 1)
for i, c1 in enumerate(s1):
current_row = [i + 1]
for j, c2 in enumerate(s2):
insertions = previous_row[j + 1] + 1
deletions = current_row[j] + 1
substitutions = previous_row[j] + (c1 != c2)
current_row.append(min(insertions, deletions, substitutions))
previous_row = current_row
return previous_row[-1]
This won't catch every spoof—determined attackers use Unicode characters that look identical to Latin letters—but it catches the common cases.

Check 4: URL Structure Anomalies

Phishing URLs tend to be longer, deeper, and messier than legitimate ones. Features that matter include:

Total URL length (phishing URLs are often 75+ characters)

Number of dots (deep subdomains)

Number of hyphens (brand-keyword combinations)

Presence of an IP address instead of a domain name

Excessive path depth (/a/b/c/d/login.html)

python
def structure_anomalies(url):
parsed = urlparse(url)
hostname = parsed.hostname or ''
flags = []
if len(url) > 75:
flags.append('long_url')
if hostname.count('.') > 3:
flags.append('deep_subdomains')
if url.count('-') > 2:
flags.append('many_hyphens')
if parsed.path.count('/') > 3:
flags.append('deep_path')
return flags
A tool called LinkSentry, which uses 111 URL features, found that directory length, URL length, and slash count were among the top five most important predictors of phishing . Structure matters.

Turning Flags into a Score (Not a Verdict)
A good detection tool doesn't say "this is phishing." It says "this URL has 4 suspicious indicators." The difference matters because every single indicator produces false positives.

Legitimate sites use "secure" and "login" in URLs. Real companies have long subdomains. Hyphens appear in plenty of safe domains.

The goal is risk scoring, not binary judgment. You might weight indicators differently:

Brand appearing only in subdomain: high weight

IP address instead of domain: high weight

Suspicious keyword + brand imitation together: elevated

Long URL alone: low weight

When multiple indicators stack up, confidence increases. A single flag means "look closer." Three or four flags means "don't enter credentials here."

The Limits of URL Analysis
Your tool won't catch everything. Sophisticated attackers use clean, short domains that look legitimate. They can register domains that contain no suspicious keywords and no brand imitation. They use HTTPS with valid certificates because certificates are free and easy to obtain .

URL analysis is one layer. It's fast, it's explainable, and it catches the majority of low-effort phishing attacks. But it should be combined with other signals: domain age (new domains are riskier), page content analysis, and visual comparison against known legitimate sites .

The most important thing your tool does is slow you down. Phishing succeeds when you click without thinking. A tool that flags suspicious patterns interrupts that automatic behavior—and that interruption is often enough to prevent the mistake.

Conclusion

Building a URL-analysis tool teaches you something most security advice glosses over: indicators are clues, not proof.

Suspicious keywords, misleading subdomains, brand imitation, and unusual URL structures all correlate with phishing. But they also appear in legitimate contexts. Your code should surface these signals, score them, and hand the decision back to you.

The value isn't in creating a perfect detector. It's in developing the habit of reading URLs critically—and understanding why a padlock alone means nothing.

Start with the checks above. Run them against URLs you encounter. See what your tool flags and ask yourself why. Over time, you'll internalize the patterns, and your first instinct when you see a strange URL will be to look, not click.

Published: October 11, 2026
← Previous Article SQL Injection Explained: How One Mistake Can Expose an Entire Database