How our 10-week Industrial Immersion sprint built an LLM-on-LLM honeypot testbed—and why it matters for defenders, researchers, and curious builders
TL;DR: We built an autonomous LLM attacker that repeatedly probed an LLM-augmented honeypot, capturing 3,055 attack sessions with 78,305 labeled commands using the MITRE ATT&CK framework. We analyzed which models keep adversaries engaged longest, which defenses deceive them better, and whether reconfiguration helps.
- GPT-4.1-mini excelled as both attacker and defender, producing the longest sessions
- MITRE coverage across all 14 tactics (Discovery dominated at ~45%)
- Reconfiguration results were mixed—some configs clearly outperformed others, but fixed intervals didn't consistently extend sessions
Why we did this
Honeypots are made to waste attacker resources and gather threat intelligence. We wanted to push the labeled datasets that can be created but human experts are scarce and standard bots are boring. LLMs can fill that gap. Our question was simple:
Can we create a dataset based on an LLM powered system and understanding how we can effect the attacks from the defenders perspective.
What we built
- Sangria (attacker): an LLM agent with tools inside a Kali container that executes commands, reasons, and self-labels each step with MITRE ATT&CK tactics (small sample validated by a second reviewer model and human).
- Blue Lagoon (honeypot): a Beelzebub/Cowrie-based decoy with optional LLM responses over SSH, configs are YAML-driven for fast swaps.
- Reconfiguration pipeline: a prompt-and-RAG workflow that proposes new honeypot configs using historical configs and vulnerability retrieval; we deploy between sessions, not mid-session.
All interactions—prompts, commands, labels, and honeypot responses—are logged into unified traces for downstream and future analysis.
How big it got
- 3,055 attack sessions, 78,305 labeled commands, across 10 trials.
- Coverage of all 14 MITRE ATT&CK tactics; Discovery ≈ 45% of actions.
- A 36-sample label spot-check showed ≈ 88.9% agreement with at least one reviewer (human or o3), with 95% CI to be between 76.6% to 96.1%.
Three findings that stood out
1. As an attacker, GPT-4.1-mini stayed in the fight longer.
In head-to-head runs, GPT-4.1-mini produced the longest average sessions (16.84 commands). o4-mini's sessions were shorter but showed the highest tactic diversity (normalized Levenshtein 0.77). GPT-4.1 was in the middle on both.
2. As a defender, GPT-4.1-mini made attackers work the hardest.
When powering the honeypot, GPT-4.1-mini yielded 35.77 commands per session on average (best), with Llama 3.3:70b next at 33.53. Detection rates varied widely: GPT-4.1-mini was detected least (63.2%), while Llama 3.3:70b was detected most (89.8%), yet still sustained long pre-detection sessions.
3. Auto-reconfiguration needs smarter triggers.
Our fixed-interval, between-session reconfigurations did not consistently lengthen sessions. Still, configurations beat the default, and one did much worse—because its prompt and host name literally said "honeypot," making detection trivial. Next step: trigger reconfigs based on novelty/plateau metrics, not a fixed counter.
How we measured things
We measured attacker efficiency through average session length, honeypot deceptiveness via detection rates and session lengths before/after discovery, and threat intelligence value using normalized Levenshtein distances between MITRE ATT&CK tactic sequences.
Session length counted bash commands until attack success, honeypot detection, or 100-command limit. Levenshtein distances compared all pairs of tactic sequences from N sessions, normalized by the longer sequence length to handle varying session durations.
Why this matters
- For defenders: LLM-backed honeypots can increase the interaction you get versus static shells, and LLM attacker can create longer, higher-quality logs to study. (See 35.77 vs. 16.08 average commands for GPT-4.1-mini vs. static.)
- For researchers: Thousands of labeled commands across all 14 tactics are ideal for training/evaluating log-labelers and sequence models.
- For builders: Our code is open-source, so you can fork it, try it out, and adapt it to your own needs.
What's next
- Adaptive reconfiguration: Switch from fixed-interval to smart triggers based on session length plateaus or behavioral novelty metrics.
- Enhanced attacker goals: Give attackers specific objectives and improve evaluation with human-in-the-loop scoring.
- Advanced analysis techniques: Explore command clustering beyond MITRE ATT&CK, next-command prediction, and reinforcement learning for optimal honeypot configurations.
Try it yourself
If you're into AI-enabled cyber-defense, dataset curation, or just want to pit your agent against ours:
- Fork the GitHub repository and follow the README
- Select the wanted attacker and defender LLM, and configure the honeypot as you like
- See how the logs from both the attacker and defender collect!
- Analyze the results using the provided scripts or your own methods