Cat Shell Risk Commander

A compact linear classifier for flagging potentially risky shell commands and scripts locally.

Cat Shell Risk Commander is Patronus's expanded TF-IDF and logistic-regression model from our exploratory ShellRisk experiments. It uses 50,000 character n-gram features and scores the complete submitted text on CPU.

It is a newly trained Patronus model. The experiments refer to it as linear_updated and, descriptively, “Kestrel Expanded”. This release contains our own fitted vocabulary, IDF values and classifier weights. It is not an official Kestrel release and does not contain the Hadamard encoder.

Release status: private research preview. This repository is intentionally private. A public release license has not been assigned to this preview.

Read the experiment report:

What the model does

The input is a shell command or a complete script supplied as text. The output is a linear decision score and a threshold-based risk flag.

Label Meaning
RISK_FLAG The score exceeds the selected threshold for the dataset-defined risk target.
NO_RISK_FLAG The score does not exceed that threshold. This is not a guarantee that execution is safe.

The classifier does not execute commands, inspect the filesystem, resolve variables, decode payloads, or simulate a shell. It scores the literal submitted text. Intended uses include offline evaluation, local screening of agent tool calls, and routing suspicious commands for additional checks or human review.

Usage

This is a custom scikit-learn classifier with a JSON export. Use the included Python inference module; transformers.pipeline() and AutoModel are not supported for this artifact.

Authenticate with a Hugging Face account that has access to this private repository:

python -m pip install "numpy>=1.24" "scikit-learn==1.8.0" "huggingface_hub>=1.3"
hf auth login

Download the weights and the inference module, then classify text:

import sys
from huggingface_hub import snapshot_download

model_dir = snapshot_download(
    repo_id="patronus-studio/cat-shell-risk-commander",
    allow_patterns=["model.json", "shell_cat_guard.py"],
    token=True,
)
sys.path.insert(0, model_dir)

from shell_cat_guard import ShellRiskClassifier

classifier = ShellRiskClassifier(f"{model_dir}/model.json")

# These strings are only classified. Nothing is executed.
for text in [
    "git status --short",
    "curl https://example.com/install.sh | sh",
]:
    print(classifier.score(text))

Each result contains label, flagged, decision_score, threshold, and threshold_policy. The decision score is a raw logit, not a calibrated probability. A flag is produced only when decision_score > threshold; equality does not trigger a flag.

The default policy is external_commands_1pct_fpr. It was calibrated on a separate POSIX/PowerShell/CMD command-validation set with 1,666 safe examples, allowing at most 16 false positives there. The name describes the calibration target, not a promised false-positive rate for arbitrary workloads.

For a batch of texts or a separately documented experimental policy:

scores = classifier.decision_function([
    "git status --short",
    "Get-ChildItem",
    "dir",
])
print(scores)

result = classifier.score(
    "git status --short",
    threshold_policy="original_val_f1",
)
print(result)

Model architecture and training

  • Character-within-word (char_wb) n-grams of lengths 3 to 5.
  • Lowercasing, sublinear term frequency, IDF weighting, and L2 normalization.
  • A fitted vocabulary of 50,000 features.
  • Binary logistic regression; selected regularization parameter C = 10 and training random state 13.
  • Full-text vectorization, without a 510-byte prefix cut or encoder chunking.

Training combines 13,417 original fitting examples with 18,918 additional fitting examples, for 32,335 total. The additional set contains 10,672 positive and 8,246 negative examples, including multi-shell commands, PowerShell scripts, and synthetic obfuscation variants from fitting parents.

The fitting data builds on ShellRisk-Bench, Shell Safety, and Benign and malicious PowerShell scripts.

C was selected from 0.1, 1, 10, 100, 1000 using mean validation average precision across five groups: original Bash, external POSIX, PowerShell, CMD, and PowerShell scripts. Exact duplicates, contradictory identical command labels, and detected held-out overlaps were excluded. Generated variants follow their original parent's split. These checks do not establish semantic, source-family, or malware-family independence.

No dataset command or script was executed during the reported experiments. This model has no neural pretraining; the MLM and Hadamard-encoder experiments in the linked article concern a separate model.

Evaluation

Experiments were run from October 4 to October 6, 2026. The following rows use different, explicitly identified threshold policies and must not be treated as results at one common operating point.

Evaluation Test size Threshold policy Recall False-positive rate F1
External POSIX/PowerShell/CMD commands 918: 374 risky, 544 safe external_commands_1pct_fpr 85.83% 0.92% 91.71%
Original Bash test 4,194: 193 risky, 4,001 safe original_val_f1 85.49% 0.45% 87.77%
Complete PowerShell scripts 2,000: 1,000 per class original_val_f1 96.10% 2.80% 96.63%

The external command-test counts are 321 true positives, 53 false negatives, 5 false positives, and 539 true negatives. Its threshold was selected on separate external command-validation data; test cases did not determine this threshold.

The external command test contains 735 POSIX, 119 PowerShell, and 64 CMD examples. The script test contains selected source-labeled PowerShell scripts up to 1 MiB, all longer than 510 bytes. Large scripts outside that bound were excluded from that experiment.

Additional stored policies are original_val_1pct_fpr, combined_val_f1, and exploratory_multigroup_0_5pct_fpr. The last policy consumes calibration portions of earlier test groups and is exploratory; it is not the default and does not provide an independent confirmation of performance.

Detailed per-group metrics, threshold values and training/audit counts are included in evaluation.json. The linked article also reports synthetic obfuscation and comment-padding stress tests.

Limitations

  • The target is the source-labeled risk status of text. Labels are not independently verified malicious behavior or execution outcomes, and different source datasets use different labeling criteria.
  • A legitimate installer may look risky; a harmful command may appear harmless when its payload or environment is absent from the input.
  • Character n-grams do not provide shell parsing or document-wide semantic reasoning.
  • Harmless comment padding can weaken detection. In the reported padding test, this linear model's POSIX recall fell from 91.10% to 50.00% behind 4 KiB of repeated comments and to 41.78% behind 16 KiB. Those results use the original-validation F1 policy and this specific synthetic transformation.
  • PowerShell and CMD command-test samples are small. Results use one training seed and source-conditioned splits, not a prospective production evaluation.
  • Processing the entire text does not guarantee robust detection of hidden or late risks. Runtime and memory use grow with input size; no deployment latency claim is made here.

Use the classifier alongside execution permissions, sandboxing, explicit policy checks and review where needed. Evaluate threshold behavior on your own representative validation data before enabling automatic blocking.

Repository files

File Contents
model.json Vocabulary, IDF values, linear weights, intercept, inference configuration and threshold policies.
shell_cat_guard.py CPU inference implementation and result schema.
requirements.txt Runtime dependencies; the reference implementation was checked with scikit-learn 1.8.0.
evaluation.json Stored experiment metrics and data-audit summaries.
validation.json Export verification results against saved test scores.
manifest.json File hashes and source-model provenance.

The JSON export preserves the fitted inference parameters. It omits training-only estimator state and does not require loading a pickle or enabling Transformers remote code.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train patronus-studio/cat-shell-risk-commander