By emmanuelygr
The Mini SIEM / Log Analyzer is a lightweight, Python-based tool designed to parse, analyze, and extract security events from Linux authentication logs (like /var/log/auth.log or /var/log/secure). It mimics the fundamental behavior of a Security Information and Event Management (SIEM) system by ingesting raw logs, matching them against predefined signatures (rules), generating security alerts, and exporting the findings.
I built this project to bridge the gap between theoretical cybersecurity concepts and practical programming skills. Understanding how large-scale SIEMs (like Splunk, QRadar, or Elastic Security) work under the hood is crucial for any security professional. This tool breaks down those complex enterprise systems into their core components: log parsing, rule matching, and alerting.
By developing and using this project, the key learning outcomes include:
- Log Analysis: Understanding the structure of standard system logs and identifying normal vs. anomalous behavior.
- Regular Expressions (Regex): Mastering regex to extract specific data points (IPs, usernames, timestamps) from unstructured text.
- SIEM Fundamentals: Gaining hands-on experience with core SIEM concepts: Parsing, Event Correlation, and Alerting.
- Python Automation: Using Python to automate repetitive security tasks (log review).
- Data Serialization: Exporting parsed data into structured formats (JSON, CSV) for further reporting.
- Regex-Based Log Parsing: Extracts meaningful data from raw log strings.
- Customizable Ruleset: Detects common events like SSH logins, failed passwords, invalid user attempts, and sudo usage.
- Brute Force Detection: Aggregates events to identify brute-force attacks (e.g., >5 failed logins from the same IP).
- Structured Output: Exports parsed events and generated alerts to both JSON and CSV formats.
- Zero Dependencies: Runs entirely on the Python Standard Library (no
pip installrequired!).
- Python 3.x: Core programming language.
re(Standard Library): For regular expression matching.json/csv(Standard Library): For exporting results.argparse(Standard Library): For handling command-line arguments.
- Logs as Evidence: Logs are the foundation of security monitoring, recording who did what, when, and where.
- SIEM (Security Information and Event Management): Systems that collect logs from across an environment, normalize them, and analyze them for threats.
- Signatures / Rules: Predefined patterns used to identify known bad behavior (e.g., "Failed password for root").
- Indicators of Compromise (IoCs): IP addresses or usernames involved in suspicious activity.
- False Positives: Alerts generated for benign activity (e.g., an administrator mistyping their password once).
- Ingestion: The script reads a log file line by line.
- Parsing (The
parse_log_filefunction): It tests each line against a dictionary of Regex rules. If a match is found, it extracts variables (likeiporuser) and creates an "Event" object. - Analysis (The
analyze_eventsfunction): It reviews the list of parsed events. It counts failed logins per IP to detect brute-force thresholds and flags attempts to log in as invalid users. - Reporting (The
export_resultsfunction): It writes the raw events and the generated alerts to JSON or CSV files for review.
mini-siem-log-analyzer/
├── analyzer.py # The main Python script
├── sample_auth.log # Mock log file for testing
├── requirements.txt # Empty (uses standard lib)
├── .gitignore # Git ignore file
├── LICENSE # MIT License
└── README.md # Project documentation
- Clone this repository or download the files.
- Ensure you have Python 3.6 or higher installed (
python --version). - Navigate to the project directory:
cd mini-siem-log-analyzer
Run the script from the command line, pointing it to a log file.
Basic Usage (defaults to JSON output):
python analyzer.py sample_auth.logExport to CSV:
python analyzer.py sample_auth.log --format csvSpecify a custom output prefix:
python analyzer.py sample_auth.log --format json --output my_investigationpython analyzer.py -hWhen running python analyzer.py sample_auth.log, the script will parse the provided mock log.
Console Output:
[*] Starting Mini SIEM Log Analysis on 'sample_auth.log'...
[*] Parsing logs...
[+] Found 14 interesting events.
[*] Analyzing events for security threats...
[+] Generated 3 security alerts.
--- SECURITY ALERTS ---
[HIGH] Brute Force Attack Detected: IP 10.0.0.5 has exceeded the threshold of 5 failed login attempts.
[MEDIUM] Invalid User Login Attempt: Attempted login with non-existent user: admin from IP 10.0.0.6
[MEDIUM] Invalid User Login Attempt: Attempted login with non-existent user: test from IP 172.16.0.5
-----------------------
[*] Exporting results...
Results exported to analysis_events.json and analysis_alerts.json
[*] Analysis complete.
It will generate two files:
analysis_events.json: Contains structured data for every matched log line.analysis_alerts.json: Contains the high-level security alerts (the brute-force detection and invalid users).
Want to improve this project? Try these exercises:
- Add a New Rule: Modify the
RULESdictionary inanalyzer.pyto detect a new event type, such as "session opened for user root". You will need to write the appropriate regex. - Geolocate IPs: Integrate an external API (like
ip-api.comoripinfo.io) to look up the geographic location of attacking IP addresses. - Real-time Monitoring: Modify the script to use
tail -flogic, continuously reading the log file as new lines are written, rather than processing it once and exiting. - Database Integration: Instead of exporting to JSON/CSV, use SQLite (
import sqlite3) to store events and alerts in a database.
- Log Tampering: Attackers often try to delete or modify logs to hide their tracks. In a real environment, logs should be forwarded to a secure, centralized server immediately.
- Sensitive Data: Logs can contain sensitive information. Ensure that any exported data complies with your organization's data handling policies.
This tool is designed for educational purposes, defensive analysis, and analyzing logs on systems you own or have explicit permission to monitor. Do not use this tool on systems or networks without authorization.
- Format Specificity: The regex rules are hardcoded for standard Linux
auth.logformats. They will fail if the log format changes (e.g., using a different syslog daemon or OS). - Performance: Loading the entire parsed dataset into memory (the
eventslist) is fine for small files but will cause memory issues on massive enterprise log files (gigabytes in size). - No Alert Routing: True SIEMs route alerts via email, Slack, or ticketing systems (Jira, ServiceNow). This script only prints to the console and writes to a file.
Q: The script runs but finds 0 events. A: The regex rules might not match the format of your specific log file. Check a sample line from your log against the regex using a site like regex101.com.
Q: Can I use this on a Windows system?
A: Yes, the Python script runs on Windows. However, Windows Event Logs (.evtx) use a completely different format than Linux text logs. You would need to export the Windows logs to a text format or use a Python library designed for .evtx files, and write entirely new Regex rules.
Building this project reinforced my understanding of string manipulation, data structures, and the logic required to correlate disparate events into a cohesive security alert. It highlighted the challenges of parsing unstructured data and the importance of having standard log formats.
- Modular Rule System: Move the rules out of the Python script and into an external configuration file (like YAML or JSON) so users can add rules without touching the code.
- Support for Other Log Types: Add parsing logic for web server logs (Apache/Nginx access logs) or firewall logs.
- Dashboarding: Create a simple Flask or Streamlit web app to visualize the JSON output (e.g., a pie chart of alert types, a bar chart of top attacking IPs).
- Anomaly Detection: Implement basic statistical analysis (e.g., "This user normally logs in at 9 AM, but just logged in at 3 AM") rather than just signature matching.
- Multi-Threading: Implement multi-threading or multi-processing to analyze multiple log files concurrently for faster processing.