LogSentry-LM / README.md
sankath's picture
Upload README.md with huggingface_hub
e51178b verified
|
Raw
History Blame Contribute Delete
7.53 kB
---
license: mit
library_name: pytorch
tags:
- log-analysis
- anomaly-detection
- security
- transformer
- gpt
- language-model
- from-scratch
pipeline_tag: text-generation
---
# LogSentry-LM
A small, from-scratch **GPT-style language model** trained on real server logs to
(1) autoregressively complete log lines and (2) **detect anomalous log lines** by
how "surprised" the model is by them (perplexity-based scoring).
It is a compact, educational decoder-only Transformer (~2M parameters) built with
plain PyTorch β€” no `transformers` modeling class β€” intended for learning and for
lightweight log anomaly detection.
> πŸš€ **Try the live demo:** https://log-sentry-lm.streamlit.app/ β€” paste log lines
> and see them scored in real time.
## Model details
- **Architecture:** decoder-only Transformer (GPT-style), causal self-attention,
pre-norm residual blocks, learned token + positional embeddings.
- **Size:** `d_model=128`, `n_heads=4`, `n_layers=3`, `max_seq_len=128` (~2.03M params).
- **Tokenizer:** Byte-Pair Encoding (`tokenizers`), vocab β‰ˆ 5.5k.
- **Training data:** combined **Loghub** samples β€” Linux + Apache + OpenSSH
(`*_2k.log`, ~6k lines total).
- **Objective:** next-token prediction (cross-entropy).
- **Framework:** PyTorch (trained on Apple Silicon MPS).
## Files in this repo
| File | Purpose |
| --- | --- |
| `logsentry_lm.pt` | Trained weights + anomaly threshold + config |
| `log_tokenizer.json` | BPE tokenizer |
| `architecture.py` | Transformer building blocks |
| `dataset.py` | Tokenizer loader + dataset |
| `train_and_generate.py` | Model class, training, generation, anomaly scoring |
| `detect_anomalies.py` | Standalone anomaly predictor |
## Usage
> ⚠️ This is a **custom architecture**, so `AutoModel.from_pretrained` does **not**
> work. Download the files and load manually:
```python
from huggingface_hub import hf_hub_download
import torch
from train_and_generate import LogSentryLM, score_line
from dataset import build_or_load_tokenizer
repo = "sankath/LogSentry-LM"
ckpt_path = hf_hub_download(repo, "logsentry_lm.pt")
tok_path = hf_hub_download(repo, "log_tokenizer.json")
ckpt = torch.load(ckpt_path, map_location="cpu")
model = LogSentryLM(vocab_size=ckpt["vocab_size"], max_seq_len=ckpt["max_seq_len"])
model.load_state_dict(ckpt["model_state"]); model.eval()
tokenizer = build_or_load_tokenizer(None, tok_path) # loads the cached tokenizer
loss, ppl = score_line(model, tokenizer, "Failed password for root from 1.2.3.4 port 22 ssh2", "cpu")
print("anomaly" if loss > ckpt["threshold"] else "normal", f"(perplexity={ppl:.1f})")
```
## How anomaly detection works
The model learns what *normal* logs look like. A line it cannot predict scores a
high per-token loss (high perplexity) and is flagged as anomalous. The cutoff is
`threshold = mean + 2Β·std` of the training corpus line scores, stored in the
checkpoint. Reference behavior:
| Input | Perplexity | Verdict |
| --- | --- | --- |
| Normal Apache `[notice]` line | ~2 | normal |
| Real SSH login attempt | ~12–16 | borderline |
| Random gibberish / injected text | >100,000 | anomaly |
## Applications & use cases
Because the model learns the "grammar" of normal system logs, it can support a
range of log-analytics and security tasks:
- **Security monitoring / intrusion detection** β€” flag suspicious log lines
(brute-force SSH attempts, injected commands, malformed requests) that don't
match learned normal patterns. This is the core "sentry" use case.
- **Operational anomaly detection (AIOps)** β€” surface rare error conditions,
misconfigurations, or failing components before they escalate, by ranking log
lines by how "surprising" they are.
- **Log triage & noise reduction** β€” instead of reading thousands of lines,
engineers review only the top-scoring (most unusual) ones.
- **Log autocompletion / templating** β€” suggest or complete log-line formats given
a prefix (useful in tooling and for synthetic log generation for testing).
- **Data-quality / format validation** β€” detect corrupted, truncated, or
non-conforming log entries in an ingestion pipeline.
- **Education** β€” a compact, readable reference for how a GPT-style Transformer,
tokenizer, and perplexity-based anomaly scoring work end to end.
### How it can be used
- **Batch scan a log file** β€” score every line and export the anomalies:
`python detect_anomalies.py --file /var/log/auth.log`
- **Score individual lines** programmatically via `score_line(...)` and compare
against the saved `threshold` (see the usage snippet above).
- **Stream integration** β€” call the scorer inside a log shipper (Fluent Bit /
Logstash / a Kafka consumer) to tag or route high-perplexity lines in near real
time.
- **Alerting** β€” wire flagged lines into Slack/email/PagerDuty or a SIEM
(Splunk, Elastic) as an enrichment signal.
- **Retrain on your own logs** β€” point `LOG_FILE` at your data and run
`train_and_generate.py` to specialize the model to your environment.
## Future scope
Natural next steps to make this research prototype more capable and production-ready:
- **Per-source models & thresholds** β€” separate models for SSH, Apache, kernel,
etc., removing the single-global-threshold weakness.
- **Session / sequence-level modeling** β€” detect anomalous *sequences* of events
(e.g. login β†’ privilege escalation β†’ data access), not just single lines.
- **Larger context & scale** β€” bigger `d_model`, more layers, longer
`max_seq_len`, and training on full Loghub datasets (millions of lines).
- **Template-aware tokenization** β€” parse logs into templates + variables (Γ  la
Drain) so IPs/IDs don't inflate the vocabulary or the anomaly score.
- **Calibrated scoring** β€” convert raw perplexity into probabilities/severity
levels, with per-line-position attribution ("which token was surprising?").
- **Real-time serving** β€” expose the scorer as a REST/gRPC service or a Hugging
Face **Space** (Gradio demo) for interactive use.
- **`transformers` integration** β€” wrap as a `PreTrainedModel` so it loads via
`AutoModel.from_pretrained` and plugs into the standard ecosystem.
- **Evaluation & benchmarking** β€” measure precision/recall on labeled anomaly
datasets (e.g. BGL, HDFS) against baselines like DeepLog and LogBERT.
- **Explainability & feedback loop** β€” let analysts confirm/dismiss flags to
continuously refine the threshold and reduce false positives.
## Limitations
- Trained on a **small, mixed** corpus (three log sources). A single global
threshold is imperfect across sources; per-source models/thresholds work better.
- Not production hardened β€” intended for education and experimentation.
- Small context window (128 tokens) and small vocab; it models line-level, not
session-level, structure.
## Attribution & licensing
- Training data derived from the **[Loghub](https://github.com/logpai/loghub)**
collection (Linux, Apache, OpenSSH samples). Please cite Loghub and consult its
terms for any redistribution of the underlying logs.
- Model code released under the MIT license.
```bibtex
@inproceedings{zhu2023loghub,
title = {Loghub: A Large Collection of System Log
Datasets for AI-driven Log Analytics},
author = {Zhu, Jieming and He, Shilin and He, Pinjia and
Liu, Jinyang and Lyu, Michael R.},
booktitle = {IEEE International Symposium on Software
Reliability Engineering (ISSRE)},
year = {2023}
}
```