Abstract
Large language model (LLM) backends increasingly generate SSH-honeypot output, yet the literature fixes one backend per system and judges realism by human evaluators, never asking which LLM is fit to play the shell or measuring the catastrophic failure of a backend leaking its own instructions. We fix one hardened unprivileged user scaffold and prompt, vary only the backend across eleven LLMs, and replace the human judge with an objective, prompt-anchored leakage metric. Across a controlled 42-command battery (20 trials each; 8736 responses) and a live adaptive corpus (8626 responses), we score six dimensions: instruction leakage, fidelity, hallucination, latency, verbosity, and stability. Exactly one backend (gemma-4-31b-it) reproduces verbatim leakage in both datasets and is disqualified; the other ten never leak. A deterministic handler layer serves 33 of 42 commands identically across backends, confining model risk to nine generative commands, where flag fabrication ranges 0–95%, latency spans an order of magnitude, and a held-out classifier identifies the backend from one response at 71% versus 9% chance, a fingerprinting risk. We report a per-dimension scorecard rather than a weight-sensitive ranking. Honeypot safety on the privilege boundary is a property of the architecture; realism, speed, cost, and one catastrophic leak are properties of the model.
IPC Classification
Keywords
€ 4.00