Primordial OS · Defensive Architecture

HIR × OAM Defense Map
Against Black-Hat LLMs

A bounded systems map for defending people, institutions, and AI runtimes from malicious LLM use, degraded automation, prompt abuse, synthetic social engineering, and agency capture.

Boundary: This artifact is defensive only. It does not teach attack construction, evasion, credential theft, malware creation, fraud, or bypass methods. It maps risk, governance, detection, and recovery through HIR and OAM.

01

Core Thesis

Black-hat LLMs are not just “bad chatbots.” They are pressure multipliers. They compress skill gaps, automate deception, scale manipulation, and remove friction from harmful intent. HIR defines what must be preserved. OAM diagnoses how agency is captured and degraded.

HIR = Defense Kernel

Honesty, Integrity, and Respect act as the governing baseline. Every request, output, memory, tool call, and escalation is checked against truthfulness, structural consistency, and human/system boundaries.

PreserveBaseline

OAM = Fault Detector

Outsourced Agency Model identifies when a system is being used to displace human judgment, intensify coercion, hide accountability, or convert people into manipulable targets.

DetectDegradation

Defense = Runtime Discipline

The goal is not one static rule. The goal is layered runtime discipline: classify, constrain, verify, isolate, log, escalate, and repair without losing respect for legitimate users.

BoundRecover
02

Threat Map

These are defensive categories, not instructions. Each category names a risk surface that an HIR/OAM runtime should recognize and contain.

Social Engineering

Automated persuasion, impersonation, synthetic intimacy, authority mimicry, and emotional pressure designed to get people to act against their interest.

Respect FailureAgency Capture

Fraud Amplification

Scams, forged context, fake support, fake invoices, counterfeit authority, and high-volume manipulation pipelines.

Honesty FailureFalse Context

Cyber Abuse Assistance

Requests that try to turn the model into a planning, coding, troubleshooting, or operational assistant for unauthorized access or harm.

Integrity BreachTool Misuse

Disinformation Engines

Mass-produced false narratives, context collapse, fake evidence, synthetic consensus, and audience-targeted manipulation.

Truth CollapseField Degradation

Jailbreak Brokerage

Attempts to route around safety boundaries, extract forbidden behavior, or convert the assistant into a policy-evasion system.

Boundary ProbeRuntime Attack

Data Exfiltration

Efforts to reveal secrets, credentials, private data, prompts, logs, or hidden operational context.

Provenance BreachTrust Theft

Automation Swarms

Use of many agents, accounts, or generated variants to overwhelm moderation, customer support, public discourse, or security review.

Scale PressureSignal Flood

Human Targeting

Using models to identify, pressure, shame, groom, extort, isolate, or exploit vulnerable people.

Respect ZeroLife-First Violation
03

HIR Defense Translation

HIR Term
Security Meaning
Defensive Question
Honesty

Reality contact, source clarity, identity truthfulness, evidence boundary.

Truth gate

Detect false premise, impersonation, fabricated authority, hidden intent, and unverifiable claims.

Is the request/output anchored to truthful context, or is it manufacturing false reality?
Integrity

Structural consistency, role fidelity, auditability, non-corruption under pressure.

Runtime gate

Preserve policy, tool boundaries, provenance, logging, and reproducible reasoning.

Does the action preserve system structure, or does it exploit a contradiction?
Respect

Dignity, consent, boundaries, consequence awareness, life-first constraint.

Human gate

Block coercion, targeting, exploitation, manipulation, and non-consensual harm.

Does this preserve human agency and boundaries, or does it convert a person into an object?
04

OAM as Degradation Detector

Black-hat LLM behavior often follows an OAM pattern: agency is moved away from a responsible person and into an automated, deniable, scaled system.

OAM Signal

Outsourced Agency

The user attempts to make the model carry intent, judgment, planning, credibility, or consequence while hiding the real actor behind automation.

DeniabilityAutomationPressure
Defense Response

Re-anchor Responsibility

Classify intent, demand lawful/authorized context, constrain outputs, protect targets, preserve logs, and redirect toward safe defensive or educational alternatives.

AuditLimitRepair
05

HIR Runtime Defense Pipeline

A black-hat resistant system should not rely on a single refusal. It should route every risky request through a layered defensive pipeline.

01
Intent IntakeClassify user goal, actor role, authority, scope, and likely target.
02
HIR GateCheck truthfulness, structural consistency, and respect for boundaries.
03
OAM ScanDetect agency outsourcing, coercion, manipulation, scaling, and deniability.
04
Tool FirewallLimit or block actions that could enable unauthorized access, fraud, or harm.
05
Safe TransformConvert unsafe requests into defensive education, policy, awareness, or repair steps.
06
Audit & LearnLog risk signals, preserve provenance, and strengthen future detection.
06

Interactive Risk Gate

Move the sliders to see how an HIR/OAM defense gate can classify a request. Lower HIR and higher OAM means higher risk.

GREEN · Defensive / Allowed

HIR is stable and OAM pressure is low. The system can answer normally while preserving boundaries.

07

Defensive Controls

For AI Runtime Designers

Build HIR gates into request classification, memory retrieval, tool invocation, user-state tracking, and escalation logic. Treat Respect as a hard safety boundary, not a style preference.

PolicyMemoryTools

For Security Teams

Monitor for synthetic social engineering, high-volume narrative generation, credential-seeking patterns, impersonation attempts, automation swarms, and tool misuse requests.

SOCDetectionTriage

For Institutions

Protect staff and customers from AI-amplified scams by verifying identity, reducing pressure-based decisions, and making safe reporting channels easy to access.

GovernanceTrainingProvenance

For Public Education

Teach people that persuasive text, synthetic authority, and emotional urgency are not proof. Slow down, verify, and preserve agency before acting.

AgencyVerificationDignity
Final invariant: A black-hat LLM attack succeeds when it breaks truth contact, corrupts decision structure, or converts people into targets. HIR defends the baseline. OAM exposes the degradation path.