Technology

The Future of Technology Starts Here

AI · Web3 · Cloud · Cyber · next-gen dev

Cyber Security

Build a Mini Intrusion Detection System With Python

Description

In the modern digital landscape, a server is bombarded with thousands of requests every single hour. While most of these are legitimate users looking for information, a concerning percentage are bots, scanners, and malicious actors probing for weaknesses. While enterprise-level Security Information and Event Management (SIEM) systems cost thousands of dollars, the core logic behind them is surprisingly accessible.

Introduction

Imagine a security guard standing at the door of a nightclub. They have a list of rules: "Don't let in anyone with a mask," or "If someone tries the door handle ten times in a row, call the police." This is essentially what a rule-based Intrusion Detection System does for a computer network.

When you build an IDS, you are essentially teaching a computer to read a diary—the server log—and look for entries that sound angry, repetitive, or strange. Server logs (like those from Nginx or Apache) record every request made to your server. They are a goldmine of information, but they are also incredibly noisy. A human cannot manually read through 50,000 lines of text to find the one IP address that is trying to break in.

That is where Python comes in. Python is the ideal language for this task because of its powerful text-processing capabilities and easy-to-read syntax. It allows us to write "rules" that filter this noise.

In this article, we will build a script that ingests a sample log, applies logic to identify specific attack signatures, and outputs a report. This isn't just a coding exercise; it is a fundamental lesson in how security operations centers (SOCs) function. Let’s start building.

Content

Setting the Stage: The Sample Data
To build our detector, we need data. We will simulate a typical web server log. Real logs can be complex, but we will use a simplified format for clarity:

[TIMESTAMP] [IP_ADDRESS] [METHOD] [STATUS_CODE] [URL]

Here is a sample of what our server.log file might look like:

text
2023-10-27 10:00:01 192.168.1.5 GET /index.html 200
2023-10-27 10:00:05 192.168.1.5 GET /about.html 200
2023-10-27 10:01:10 203.0.113.42 POST /login 401
2023-10-27 10:01:12 203.0.113.42 POST /login 401
2023-10-27 10:01:14 203.0.113.42 POST /login 401
2023-10-27 10:02:00 198.51.100.7 GET /admin 403
The Logic: Rules of Engagement
Our Mini IDS will operate on a simple "If This, Then That" basis. We will define three primary rules:

Brute Force Detection: If a single IP address accumulates more than 3 failed login attempts (Status 401) within a short window, flag it.

Suspicious Status Codes: If an IP hits a high number of "403 Forbidden" or "404 Not Found" errors, they might be scanning for vulnerabilities.

Unusual Request Patterns: If an IP makes an abnormally high number of requests in a single second (simulating a DoS attack or aggressive bot).

The Code
Let's write the Python code. We will use a dictionary to store the activity of each IP address.

python
import re
from collections import defaultdict

# Configuration Rules
FAILED_LOGIN_THRESHOLD = 3
REQUEST_THRESHOLD = 5
SUSPICIOUS_CODES = ['403', '404']

# Sample Log Data (In a real scenario, read from a file)
log_data = """
2023-10-27 10:01:10 203.0.113.42 POST /login 401
2023-10-27 10:01:12 203.0.113.42 POST /login 401
2023-10-27 10:01:14 203.0.113.42 POST /login 401
2023-10-27 10:01:15 203.0.113.42 POST /login 401
2023-10-27 10:02:00 198.51.100.7 GET /admin 403
2023-10-27 10:02:01 198.51.100.7 GET /wp-admin 403
2023-10-27 10:02:02 198.51.100.7 GET /phpmyadmin 403
2023-10-27 10:02:03 198.51.100.7 GET /config 403
2023-10-27 10:03:00 192.168.1.5 GET /index.html 200
2023-10-27 10:03:01 192.168.1.5 GET /about.html 200
"""

def parse_log_line(line):
# Regex to capture timestamp (simplified), IP, Method, URL, and Status
pattern = r"(\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2}) (\d{1,3}\.\d{1,3}\.\d{1,3}\.\d{1,3}) (\w+) (/\S+) (\d{3})"
match = re.search(pattern, line)
if match:
return match.groups()
return None

def analyze_logs(log_content):
ip_stats = defaultdict(lambda: {'failed_logins': 0, 'suspicious_hits': 0, 'total_requests': 0})
alerts = []

for line in log_content.strip().split('\n'):
parsed = parse_log_line(line)
if not parsed:
continue

timestamp, ip, method, url, status = parsed

# Update stats
ip_stats[ip]['total_requests'] += 1

if status == '401':
ip_stats[ip]['failed_logins'] += 1

if status in SUSPICIOUS_CODES:
ip_stats[ip]['suspicious_hits'] += 1

# Apply Rules
for ip, stats in ip_stats.items():
if stats['failed_logins'] >= FAILED_LOGIN_THRESHOLD:
alerts.append(f"[ALERT] Brute Force Detected: {ip} has {stats['failed_logins']} failed logins.")

if stats['suspicious_hits'] >= 3: # Arbitrary threshold for scanning
alerts.append(f"[ALERT] Vulnerability Scanning: {ip} triggered {stats['suspicious_hits']} suspicious errors.")

if stats['total_requests'] > REQUEST_THRESHOLD:
alerts.append(f"[WARNING] High Traffic: {ip} made {stats['total_requests']} requests.")

return alerts

# Run the detector
alerts = analyze_logs(log_data)
print("--- IDS REPORT ---")
for alert in alerts:
print(alert)
The Output
When you run this script, the output will be:

text
--- IDS REPORT ---
[ALERT] Brute Force Detected: 203.0.113.42 has 4 failed logins.
[ALERT] Vulnerability Scanning: 198.51.100.7 triggered 4 suspicious errors.
The Elephant in the Room: False Positives
In the code above, we set a threshold of 3 failed logins. But what if a user simply forgot their password and tried four times? Our IDS would flag them as a hacker. This is a False Positive.

In the real world, false positives are the bane of a security analyst's existence. If your IDS screams "HACKER!" every time someone types their password wrong, analysts will eventually ignore the alerts—a phenomenon known as "alert fatigue."

To reduce false positives, we could:

Increase thresholds: Maybe 10 failed logins is more realistic than 3.

Whitelisting: Exclude internal IP addresses (like 192.168.1.5 in our example) from brute force rules.

Time windows: A brute force attack happens in seconds. A user forgetting a password happens over minutes. We could add logic to check the time difference between attempts.

Where This System Falls Short
Our Mini IDS is "signature-based." It looks for specific known bad behaviors. It cannot detect a "Zero Day" attack (a brand new exploit) because it doesn't know what to look for. It also cannot detect a slow, low-and-slow attack where the hacker tries one password per day to avoid triggering the threshold.

Conclusion

Building a Mini Intrusion Detection System with Python is a fantastic way to demystify the world of cybersecurity. We have successfully created a script that parses logs, applies logical rules, and identifies potential threats like brute force attacks and vulnerability scanning.

However, the most important lesson from this project is understanding the limitations of simple rule-based detection. The code we wrote is a starting point, not a finish line. It highlights the constant battle between security and usability. If you make the rules too strict, you block legitimate users (false positives). If you make them too loose, you miss the attackers (false negatives).

To take this project further, consider adding features like:

Email Notifications: Have the script send you an email when an alert is triggered.

Geo-IP Lookup: Integrate a library to find the geographical location of the attacker's IP.

Database Storage: Instead of printing to the console, store the logs in an SQLite database for long-term analysis.

Cybersecurity is a cat-and-mouse game. By understanding how the "cats" (detectors) work, you become a better defender. Keep coding, keep analyzing, and stay secure.

Published: October 11, 2026
← Previous Article Can AI Write Secure Code? Testing AI-Generated Code for Vulnerabilities