PhishGuard:基于 Google Gemma 2 的零泄露鱼叉式钓鱼威胁扫描器
PhishGuard: Zero-Leakage Spear-Phishing Threat Scanner Powered by Google Gemma 2
PhishGuard 是一款基于 Google Gemma 2 9B-IT 的零泄露鱼叉式钓鱼威胁扫描器,采用本地隐私屏蔽、11 维启发式引擎与 Gemma 2 推理的四阶段流水线,代码以 Apache 2.0 协议开源。它在上传前本地脱敏信用卡、学号、马来西亚 NRIC、电话与邮箱,并将钓鱼链接转为 hxxps 安全字符串,输出含 SHA-256 证据链的可下载审计报告。
This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
I built PhishGuard — a zero-leakage, forensic-grade spear-phishing threat scanner and incident response auditor designed to protect everyday users from sophisticated social engineering without surrendering their personal data to public cloud servers.
The Friend & The Problem
My close friend and university classmate, Alex, almost fell for an alarming spear-phishing email.
The attacker had obtained leaked enrollment records: the email cited Alex's exact student matriculation ID, mentioned his outstanding semester balance, and included an urgent link demanding he "verify identity within 12 hours" to unlock an emergency bursary disbursement. It was emotionally coercive and terrifyingly convincing.
When I urged Alex to analyze the email using AI, he froze:
"Can I actually paste this into ChatGPT or an online scanner? It has my real student ID, private balance, and reference codes in it..."
That moment exposed the fundamental paradox of modern cybersecurity tooling:
The messages most in need of adversarial threat analysis are precisely the ones least safe to upload.
Public AI chatbots store prompts for training telemetry, while commercial security gateways index submitted payloads. For a student or privacy-conscious friend, using existing tools to verify a threat creates a brand new data-leakage risk.
PhishGuard was built specifically for Alex — and anyone who needs military-grade threat analysis with zero data exposure.
Demo
PhishGuard is deployed across dual live instances to ensure seamless evaluation without cold-start interruptions:
- 🚀 Official Cloud Deployment (Render Blueprint): https://phishguard-gemma.onrender.com/ (Primary submission for Best Use of Render)
- ⚡ Instant Mirror (Zero-Wait Streamlit Cloud): https://phishguard-gemma.streamlit.app/ (No cold-start fallback)
- 🔒 100% Air-Gapped Local Mode: Clone and run with zero external internet telemetry needed.

Figure 1: Real-time forensic threat scoring and calibrated risk gauge identifying a high-confidence spear-phishing attack.

Figure 2: PhishGuard's Privacy Shield automatically redacting sensitive student IDs and bank card numbers locally before any telemetry is generated.
Key Capabilities Walkthrough
- Local Privacy Shield (Pre-Audit PII Redaction): Before any reasoning occurs, credit cards (Luhn format), student IDs, Malaysian NRIC, phone numbers, and emails are automatically scrubbed locally.
-
Defanged URL Inspector: Embedded phishing links are automatically converted to safe forensic strings (
hxxps://...[.]...) to prevent accidental clicks. - Multi-Vector Forensic Assessment: Combines an 11-dimensional heuristic engine with Google Gemma 2 9B instruction-tuning.
-
Downloadable Incident Response Report: Generates a formal, printable
.txtaudit report complete with SHA-256 evidence chain verification for enterprise or campus IT submission. - Adaptive Industrial Dual-Theme: Seamlessly synchronizes between Obsidian Dark and Titanium Light precision modes via native settings.
Code
The complete source code is open-sourced under the Apache 2.0 License:
Hao610
/
phishguard-gemma
A privacy-first, air-gapped threat scanner against spear-phishing & social engineering. Built for a friend, powered by Google Gemma 2.
🛡️ PhishGuard
Zero-Leakage Spear-Phishing Threat Scanner & Forensic Auditor
Built for a Friend (Alex) | Hacktoberfest 2026 Challenge
🌐 Live Instances:
- 🚀 Official Render Cloud: phishguard-gemma.onrender.com (Official Hackathon Host)
- ⚡ Instant Streamlit Mirror: phishguard-gemma.streamlit.app (Zero Cold-Start)
💡 The Origin: Built for Alex
Alex almost fell for a spear-phishing email that quoted his real student ID and tuition balance.
He wanted to verify it with AI, but stopped:
"Can I paste this email into ChatGPT? It has my real student ID and financial record in it..."
That moment exposed the core paradox of modern anti-phishing tools:
The messages most in need of analysis are the ones least safe to upload.
PhishGuard was built to solve this. It inverts the paradigm by running a privacy-first, dual-stage pipeline:
- Local Privacy Shield: Automatically strips Personal Identifiable Information (PII) before any telemetry leaves your machine.
- Offline-First Heuristics + Gemma 2 Intelligence: Detects…
- GitHub Repository: https://github.com/Hao610/phishguard-gemma
- Automated Test Suite: 13/13 passing security unit tests verifying homoglyph detection, Luhn validation, IP URLs, and header parsing.
-
Infrastructure Blueprint: Included
render.yamlfor one-click Infrastructure-as-Code deployment.
How I Built It
PhishGuard employs a multi-stage hybrid defense pipeline centered on Google Gemma 2:
┌─────────────────────────┐
│ Incoming Raw Message │
└────────────┬────────────┘
│
▼
┌─────────────────────────┐
│ Stage 1: Privacy Shield│
│ (Local PII Redaction) │
└────────────┬────────────┘
│ Sanitized Text
┌──────────────────┴──────────────────┐
▼ ▼
┌───────────────────────────┐ ┌───────────────────────────┐
│ Stage 2: Heuristic Core │ │ Stage 3: Gemma 2 Engine │
│ - 11 Threat Vectors │ │ - Google Gemma 2 9B-IT │
│ - Weighted Urgency Regex │ │ - Structured JSON Output │
│ - Homoglyph Detection │ │ - Adversarial Prompting │
└─────────────┬─────────────┘ └─────────────┬─────────────┘
│ │
└──────────────────┬──────────────────┘
│
▼
┌─────────────────────────┐
│ Stage 4: SOC Report Gen │
│ - Calibrated Risk Score │
│ - SHA-256 Evidence Chain│
│ - Defanged IoC Artifacts│
└─────────────────────────┘
1. The Open-Weight AI Architecture (Google Gemma 2)
PhishGuard leverages Gemma 2 9B-IT using its native turn-based tokenization format:
<start_of_turn>user
You are PhishGuard, an expert cybersecurity forensic auditor...
[Sanitized Message & Pre-extracted Heuristic Indicators]
<end_of_turn>
<start_of_turn>model
Gemma 2 was chosen specifically because of its state-of-the-art reasoning density relative to parameter size. Unlike general chat models that offer conversational advice, Gemma 2 was constrained with a strict schema to act as a Security Operations Center (SOC) Tier-2 Analyst, outputting calibrated risk scores, identified deception vectors, and actionable non-technical remediation plans for Alex.
2. The 11-Dimensional Threat Heuristic Engine
To ensure Gemma 2 has rich adversarial context (and to provide a 100% offline fallback when operating in air-gapped environments), a local static engine analyzes:
- Weighted Urgency Velocity: Weighted regular expressions detecting artificial panic windows ("within 12 hours", "account locked", "legal escalation").
-
Lookalike / Homoglyph Impersonation: RegEx traps identifying deceptive domains (e.g.,
wellsfarg0.com,paypa1.com). -
Infrastructure Anomalies: Raw IP address destinations (
http://192.168...), suspicious cheap TLDs (.xyz,.top,.tk), and free-hosting phisher staging sites (pages.dev,netlify.app).
3. Deploying with Render Blueprint
PhishGuard ships with an Infrastructure-as-Code render.yaml specification configured for Render's Python runtime. It builds automatically upon commit, pins dependency integrity, and orchestrates Streamlit with headless flags to guarantee high uptime.
Why Does Open Innovation Matter?
In security and privacy tooling, open innovation is not just a feature — it is an existential requirement.
1. Privacy Cannot Rely on Closed Promises
Closed-source commercial AI providers operate as black boxes. When a user pastes a spear-phishing email containing personal tuition records, health notices, or payroll details, they are forced to trust a vendor's Terms of Service. Open weights allow users to audit the exact model architecture, run inference locally with tools like Ollama or local transformers, and guarantee mathematically that zero packets leave their network.
2. Preventing Attack Surface Expansion
During our security benchmarking tests, we observed a crucial phenomenon: When an AI model deliberates over untrusted, adversary-crafted inputs in a cloud environment, that external pipeline itself becomes an expanded attack surface. Open innovation enables us to build a layered defense — stripping PII and defanging payloads locally before any cognitive evaluation takes place.
3. Democratizing Enterprise-Grade Defense
Sophisticated spear-phishing used to require enterprise-grade SOC tooling (like Proofpoint or Darktrace) costing tens of thousands of dollars. Open models like Google Gemma 2 bring that same level of deep semantic social-engineering detection to ordinary individuals like Alex for zero dollars.
Prize Categories
I am officially entering PhishGuard into the following categories:
- 🏆 Best Use of Gemma: Engineered natively around the Google Gemma 2 open architecture, implementing adversarial turn-based prompting, schema-constrained forensic reasoning, and privacy-preserving pre-scrubbing.
- 🚀 Best Use of Render: Fully automated Infrastructure-as-Code deployment utilizing an official
render.yamlblueprint, hosted live on Render's Web Service platform. - 🌟 Overall Winner ($250): A comprehensive, real-world open-source solution solving an authentic human problem with production-grade engineering, full unit-test coverage, and exemplary documentation.
Built with ❤️ for Alex and the open-source community.
来源:Google AI:DEV 作者专属(RSS) · dev.to