We Don't Need a Large Language Model for That: When Traditional Machine Learning and Encoders Beat Decoder LLMs in Security Classification
Large language models (LLMs) have become an attractive option for cybersecurity classification tasks because they are easy to apply to text-like security data and require little task-specific feature engineering. However, LLM inference can be expensive, general-purpose LLM architectures are not optimized for structured classification, and their effectiveness relative to traditional machine learning remains unclear for well-labeled security datasets. We compare LLMs and traditional machine learning models on two common cybersecurity classification tasks: phishing email detection and network attack classification. Our evaluation includes standalone zero-shot LLM classifiers, XGBoost, frozen and fine-tuned encoder models including ModernBERT, and two hybrid ML-LLM ensemble designs: a router that invokes an LLM only when the ML model has low confidence, and a meta-learner that trains a second-stage model on both ML and LLM outputs. Across our evaluated datasets, no standalone decoder LLM or decoder-LLM ensemble meaningfully outperforms the strongest traditional ML baselines. On phishing email detection, zero-shot LLMs perform strongly, but embedding-based ML pipelines and fine-tuned ModernBERT match or exceed them: zero-shot Claude reaches 0.970 AUCPR, while fine-tuned ModernBERT reaches 0.998. On network attack classification, zero-shot LLMs perform at or near the class prior across multiple feature representations, while XGBoost remains strong on named and aggregated tabular features; on aggregated flow statistics, XGBoost reaches 0.845 while zero-shot LLMs sit near the 0.403 prior. Fine-tuning an encoder model recovers useful signal from serialized flow statistics, but it ties rather than beats XGBoost. We further evaluate generalization and operational cost on new captures: under campaign-disjoint splits of CTU-13 at its observed 2.5% prior, XGBoost retains its ranking advantage while two of three zero-shot decoders remain near the class prior (Gemma3 retains partial signal), and a subtype-decomposed false-positive analysis of phishing detectors reveals a bulk-spam failure mode, invisible to pooled AUCPR, that a matched-budget hard-negative training intervention largely removes.
The ensemble results reinforce this. On the network tasks, routing uncertain examples to an at-chance LLM reduces AUCPR as more traffic is escalated, and stacking decoder LLM outputs with traditional model outputs leaves flow performance unchanged. The picture is more nuanced where the LLM has signal: routing and stacking can improve weaker models, and strong feature representations can make smaller decoder models competitive with larger frontier LLMs. Still, these findings suggest that practitioners building cybersecurity classification systems with sufficiently large, well-labeled datasets should evaluate traditional ML and encoder-based classification methods before adopting decoder LLMs or hybrid LLM ensembles.