|
Download README.md from auditing/README: direct link, hf CLI and curl.
- Browser
- Download file 26.8 kB
-
https://huggingface.co/spaces/auditing/README/resolve/main/README.md
- Command line
-
hf download hf://spaces/auditing/README/README.md
-
curl -L -o README.md https://huggingface.co/spaces/auditing/README/resolve/main/README.md
26.8 kB
| # Auditing | |
| **Auditing** is a Hugging Face organization focused on the **independent examination, traceability, verification and assurance of AI systems**. | |
| The project explores how models, agents, data pipelines, tools and AI infrastructure can be audited in a way that is **evidence-based, reproducible and useful in real-world deployments**. | |
| As AI systems become more autonomous, organizations need to answer more than: | |
| > **Does the system perform well?** | |
| They also need to ask: | |
| > **What happened, why did it happen, which components were involved, which controls were applied, and can the result be independently verified?** | |
| That is the focus of **Auditing**. | |
| --- | |
| ## What is AI auditing? | |
| **AI auditing** is the structured examination of an AI system, its behavior, controls, evidence and operational context against defined technical, organizational or regulatory criteria. | |
| An AI audit may examine: | |
| - models, | |
| - prompts and system instructions, | |
| - datasets, | |
| - retrieval systems, | |
| - agent behavior, | |
| - tool calls, | |
| - permissions, | |
| - memory, | |
| - logs and traces, | |
| - evaluation results, | |
| - safety controls, | |
| - infrastructure, | |
| - human approvals, | |
| - governance processes, | |
| - incidents, | |
| - documentation, | |
| - model and system changes. | |
| The goal is not merely to collect logs. | |
| The goal is to create **verifiable evidence** about how an AI system was designed, deployed and operated. | |
| --- | |
| ## Auditing vs. Evaluation vs. Validation vs. Observability | |
| These disciplines overlap, but they are not the same. | |
| | Discipline | Core question | | |
| |---|---| | |
| | **Evaluation** | How well does the AI system perform? | | |
| | **Validation** | Is the system suitable for the intended purpose and requirements? | | |
| | **Observability** | What is happening inside the system during operation? | | |
| | **Auditing** | Can system behavior, controls and evidence be independently examined and verified? | | |
| A useful mental model is: | |
| ```text | |
| EVALUATION | |
| How good is the system? | |
| │ | |
| ▼ | |
| VALIDATION | |
| Is it suitable for the intended use? | |
| │ | |
| ▼ | |
| OBSERVABILITY | |
| What is happening in operation? | |
| │ | |
| ▼ | |
| AUDITING | |
| Can we reconstruct and verify what happened? | |
| ``` | |
| Auditing can use outputs from all three other layers. | |
| --- | |
| # The AI Audit Stack | |
| An AI audit can span the complete lifecycle of a system. | |
| ```text | |
| AI AUDIT STACK | |
| SCOPE | |
| │ | |
| ┌───────────────┼───────────────┐ | |
| │ │ │ | |
| MODEL DATA AGENT | |
| │ │ │ | |
| └───────────────┼───────────────┘ | |
| │ | |
| CONTROLS | |
| │ | |
| ┌─────────────────────┼─────────────────────┐ | |
| │ │ │ | |
| Permissions Policies Guardrails | |
| │ │ │ | |
| └─────────────────────┼─────────────────────┘ | |
| │ | |
| EVIDENCE | |
| │ | |
| ┌─────────────────────┼─────────────────────┐ | |
| │ │ │ | |
| Logs Traces Tests | |
| │ │ │ | |
| └─────────────────────┼─────────────────────┘ | |
| │ | |
| AUDIT FINDINGS | |
| │ | |
| REMEDIATION | |
| │ | |
| RE-AUDIT / REVIEW | |
| ``` | |
| A mature audit process connects technical evidence with clearly defined requirements and controls. | |
| --- | |
| # What can be audited? | |
| ## 1. Models | |
| Model auditing can examine: | |
| - model identity and provenance, | |
| - architecture, | |
| - version, | |
| - weights, | |
| - model card, | |
| - training and fine-tuning information, | |
| - benchmark claims, | |
| - evaluation results, | |
| - known limitations, | |
| - licensing, | |
| - safety behavior, | |
| - model changes over time. | |
| A key audit question is: | |
| > **Can the deployed model be traced to a specific version and documented source?** | |
| --- | |
| ## 2. AI agents | |
| Agentic systems introduce additional audit requirements because they can make decisions, call tools and execute actions. | |
| An agent audit may include: | |
| ```text | |
| Goal | |
| │ | |
| Context | |
| │ | |
| Model | |
| │ | |
| Plan | |
| │ | |
| Tool Selection | |
| │ | |
| Tool Call | |
| │ | |
| Permission Check | |
| │ | |
| Action | |
| │ | |
| Observation | |
| │ | |
| Next Decision | |
| │ | |
| Outcome | |
| ``` | |
| For every important action, an auditor may want to know: | |
| - Which agent initiated it? | |
| - Which model version was used? | |
| - What context was available? | |
| - Which tool was selected? | |
| - Which parameters were passed? | |
| - Which permissions were active? | |
| - Was human approval required? | |
| - What did the tool return? | |
| - What happened next? | |
| - Was the final result successful? | |
| - Were any safety controls triggered? | |
| This creates the foundation for an **Agent Audit Trail**. | |
| --- | |
| ## 3. Tool use | |
| Tools turn AI from a prediction system into an action system. | |
| Relevant audit dimensions include: | |
| - tool inventory, | |
| - tool ownership, | |
| - tool permissions, | |
| - argument validation, | |
| - authentication, | |
| - authorization, | |
| - execution logs, | |
| - error handling, | |
| - rate limits, | |
| - external side effects, | |
| - data exposure, | |
| - human approval. | |
| A tool call should ideally be traceable from: | |
| ```text | |
| Agent Decision | |
| │ | |
| ▼ | |
| Tool Request | |
| │ | |
| ▼ | |
| Authorization | |
| │ | |
| ▼ | |
| Execution | |
| │ | |
| ▼ | |
| Tool Response | |
| │ | |
| ▼ | |
| Agent Follow-up | |
| ``` | |
| --- | |
| ## 4. Data and retrieval | |
| AI auditing can also examine the data layer. | |
| This may include: | |
| - data provenance, | |
| - data quality, | |
| - access rights, | |
| - retention, | |
| - sensitive information, | |
| - retrieval sources, | |
| - vector databases, | |
| - document versions, | |
| - grounding evidence, | |
| - citation behavior, | |
| - dataset changes. | |
| For retrieval-augmented systems, an important question is: | |
| > **Can the generated output be traced back to the information that supported it?** | |
| --- | |
| ## 5. Prompts and system instructions | |
| Prompts can materially affect model behavior. | |
| Audit evidence may therefore include: | |
| - system prompts, | |
| - prompt templates, | |
| - version history, | |
| - hidden instructions, | |
| - dynamic context, | |
| - policy prompts, | |
| - tool instructions, | |
| - prompt changes. | |
| Prompt versioning becomes particularly important in production systems where behavior changes without changing the underlying model. | |
| --- | |
| ## 6. Memory | |
| Agent memory introduces persistent state. | |
| An audit may examine: | |
| - what information is stored, | |
| - why it was stored, | |
| - how long it is retained, | |
| - who can access it, | |
| - whether the memory can be changed, | |
| - whether the user can correct it, | |
| - whether memory affected a decision. | |
| A useful audit record can connect: | |
| ```text | |
| Stored Memory | |
| │ | |
| ▼ | |
| Retrieved Memory | |
| │ | |
| ▼ | |
| Agent Context | |
| │ | |
| ▼ | |
| Decision | |
| │ | |
| ▼ | |
| Action | |
| ``` | |
| --- | |
| ## 7. Permissions and identity | |
| Autonomous systems should not be treated as having unlimited authority. | |
| Auditing can examine: | |
| - agent identity, | |
| - service accounts, | |
| - API credentials, | |
| - role-based access, | |
| - scoped permissions, | |
| - delegation, | |
| - temporary permissions, | |
| - approval workflows, | |
| - privileged actions. | |
| A critical principle is: | |
| > **The ability of an AI system to perform an action should be distinguishable from its ability to decide that the action is appropriate.** | |
| --- | |
| ## 8. Human oversight | |
| Many systems include human review or approval. | |
| Audit evidence can capture: | |
| - when human review was required, | |
| - who approved an action, | |
| - what information the reviewer saw, | |
| - whether the recommendation was accepted, | |
| - whether the human modified the result, | |
| - escalation events, | |
| - override decisions. | |
| This creates traceability between automated and human decisions. | |
| --- | |
| # Agent Audit Trail | |
| A practical audit trail for an AI agent might contain: | |
| | Field | Example purpose | | |
| |---|---| | |
| | Agent ID | Identify the executing agent | | |
| | Session ID | Group related activity | | |
| | Model | Record model and version | | |
| | Goal | Capture intended objective | | |
| | Input | Record user/system input | | |
| | Context | Identify relevant runtime context | | |
| | Tool | Record selected tool | | |
| | Tool arguments | Record requested action | | |
| | Permission | Record authority used | | |
| | Human approval | Record required approval | | |
| | Observation | Record tool/environment response | | |
| | Decision | Record next agent step | | |
| | Output | Record final result | | |
| | Cost | Track execution economics | | |
| | Latency | Track performance | | |
| | Safety event | Record control activation | | |
| | Timestamp | Establish chronology | | |
| | Trace ID | Connect events across systems | | |
| The exact schema depends on the use case, risk level and technical architecture. | |
| --- | |
| # From logging to auditability | |
| Logging alone does not automatically create auditability. | |
| An audit-ready system needs evidence that is: | |
| **Structured** | |
| Events should have consistent fields. | |
| **Timestamped** | |
| The sequence of events should be reconstructable. | |
| **Versioned** | |
| Models, prompts, policies and tools should be identifiable. | |
| **Tamper-aware** | |
| Audit evidence should be protected against unauthorized modification. | |
| **Contextual** | |
| A log entry without relevant context may not explain what happened. | |
| **Searchable** | |
| Auditors need to find related events across systems. | |
| **Retained appropriately** | |
| Evidence must exist long enough to support the intended audit process. | |
| **Interpretable** | |
| Technical traces should be understandable to the intended reviewer. | |
| --- | |
| # AI Audit Lifecycle | |
| A practical audit lifecycle can look like this: | |
| ```text | |
| 1. DEFINE SCOPE | |
| │ | |
| ▼ | |
| 2. IDENTIFY SYSTEM | |
| │ | |
| ▼ | |
| 3. MAP RISKS & CONTROLS | |
| │ | |
| ▼ | |
| 4. COLLECT EVIDENCE | |
| │ | |
| ▼ | |
| 5. TEST CONTROLS | |
| │ | |
| ▼ | |
| 6. REPLAY / RECONSTRUCT EVENTS | |
| │ | |
| ▼ | |
| 7. DOCUMENT FINDINGS | |
| │ | |
| ▼ | |
| 8. REMEDIATE | |
| │ | |
| ▼ | |
| 9. RE-TEST | |
| │ | |
| ▼ | |
| 10. CONTINUOUS ASSURANCE | |
| ``` | |
| Not every audit requires every step in the same form. | |
| The scope should reflect the system, intended use and risk. | |
| --- | |
| # Audit scope | |
| One of the most important audit decisions is defining scope. | |
| Possible scopes include: | |
| ### Model audit | |
| Focus on a specific model and its documented properties. | |
| ### Agent audit | |
| Focus on autonomous decisions, tool use and actions. | |
| ### Application audit | |
| Focus on the complete AI-powered application. | |
| ### Data audit | |
| Focus on data sources, provenance, quality and access. | |
| ### Infrastructure audit | |
| Focus on serving, runtime, identity, permissions and deployment. | |
| ### Process audit | |
| Focus on organizational controls and lifecycle processes. | |
| ### Incident audit | |
| Reconstruct a specific failure or harmful event. | |
| ### Continuous audit | |
| Continuously collect and test evidence from production systems. | |
| --- | |
| # Evidence types | |
| AI audits can draw from many evidence sources. | |
| ```text | |
| Model cards | |
| Technical reports | |
| Licenses | |
| System prompts | |
| Configuration | |
| Source code | |
| Model versions | |
| Dataset metadata | |
| Evaluation reports | |
| Red-team results | |
| Traces | |
| Application logs | |
| Tool logs | |
| API logs | |
| Approval records | |
| Access-control records | |
| Incident reports | |
| Change history | |
| Deployment manifests | |
| Monitoring dashboards | |
| ``` | |
| The strength of an audit depends heavily on the quality of its evidence. | |
| --- | |
| # Control categories | |
| A useful audit framework may group controls into categories. | |
| ## Governance controls | |
| - ownership, | |
| - accountability, | |
| - policies, | |
| - documentation, | |
| - approval processes, | |
| - change management. | |
| ## Model controls | |
| - model selection, | |
| - versioning, | |
| - evaluation, | |
| - validation, | |
| - safety testing, | |
| - release criteria. | |
| ## Data controls | |
| - provenance, | |
| - access, | |
| - quality, | |
| - retention, | |
| - privacy, | |
| - integrity. | |
| ## Agent controls | |
| - autonomy limits, | |
| - tool permissions, | |
| - delegation, | |
| - escalation, | |
| - approval requirements. | |
| ## Security controls | |
| - authentication, | |
| - authorization, | |
| - secrets, | |
| - sandboxing, | |
| - isolation, | |
| - incident response. | |
| ## Operational controls | |
| - monitoring, | |
| - observability, | |
| - logging, | |
| - rollback, | |
| - reliability, | |
| - cost limits. | |
| ## Human controls | |
| - review, | |
| - escalation, | |
| - override, | |
| - accountability, | |
| - training. | |
| --- | |
| # Auditing autonomous AI systems | |
| Autonomous systems change the audit problem. | |
| Traditional software generally executes predefined logic. | |
| An agent may dynamically choose: | |
| - which subtask to perform, | |
| - which tool to call, | |
| - which information to retrieve, | |
| - which model to use, | |
| - whether to continue, | |
| - when to stop. | |
| This creates a new audit requirement: | |
| > **Audit the decision trajectory, not only the final output.** | |
| Two agent runs can produce the same answer while following very different paths. | |
| One path may be safe and efficient. | |
| Another may involve unnecessary data access, incorrect tool calls or excessive cost. | |
| This is why trajectory-level evidence is important. | |
| --- | |
| # Auditability by design | |
| Auditability is easier when it is designed into the system from the beginning. | |
| A useful architecture may expose: | |
| ```text | |
| AI APPLICATION | |
| │ | |
| ORCHESTRATION | |
| │ | |
| ┌─────────────┼─────────────┐ | |
| │ │ │ | |
| Model Tools Memory | |
| │ │ │ | |
| └─────────────┼─────────────┘ | |
| │ | |
| TRACING | |
| │ | |
| LOGGING | |
| │ | |
| AUDIT STORE | |
| │ | |
| ┌────────────┼────────────┐ | |
| │ │ │ | |
| Review Testing Reporting | |
| ``` | |
| Auditability should not be an afterthought added only after an incident. | |
| --- | |
| # Continuous auditing | |
| AI systems can change quickly. | |
| Changes may include: | |
| - model updates, | |
| - prompt updates, | |
| - retrieval data updates, | |
| - tool changes, | |
| - new permissions, | |
| - policy changes, | |
| - infrastructure changes, | |
| - new agent capabilities. | |
| A one-time audit therefore provides only a point-in-time view. | |
| Continuous auditing can supplement periodic reviews by automatically checking selected evidence and controls. | |
| Examples: | |
| - unapproved model changes, | |
| - new tools, | |
| - permission expansion, | |
| - missing trace data, | |
| - unusual tool usage, | |
| - failed safety checks, | |
| - excessive costs, | |
| - unexpected data access, | |
| - repeated human overrides. | |
| --- | |
| # Auditing multi-agent systems | |
| Multi-agent architectures add another layer of complexity. | |
| An audit may need to reconstruct: | |
| ```text | |
| User | |
| │ | |
| ▼ | |
| Supervisor Agent | |
| │ | |
| ├── Research Agent | |
| │ └── Web / Retrieval | |
| │ | |
| ├── Analysis Agent | |
| │ └── Models / Data | |
| │ | |
| └── Action Agent | |
| └── External Tool | |
| ``` | |
| Important questions include: | |
| - Which agent delegated the task? | |
| - Which agent had authority? | |
| - Was context transferred correctly? | |
| - Which agent produced the final decision? | |
| - Could one agent escalate permissions through another? | |
| - Can the complete chain be reconstructed? | |
| Audit evidence should preserve delegation relationships. | |
| --- | |
| # Auditing model routing | |
| Systems increasingly route tasks between different models. | |
| Auditability may require capturing: | |
| - routing decision, | |
| - routing policy, | |
| - available models, | |
| - selected model, | |
| - model version, | |
| - reason or rule for selection, | |
| - fallback behavior, | |
| - cost, | |
| - latency, | |
| - outcome. | |
| This is especially important when different models have different capabilities, licenses or deployment constraints. | |
| --- | |
| # Auditing RAG systems | |
| Retrieval-Augmented Generation adds additional audit dimensions: | |
| ```text | |
| Question | |
| │ | |
| Retrieval Query | |
| │ | |
| Retrieved Documents | |
| │ | |
| Document Versions | |
| │ | |
| Context Assembly | |
| │ | |
| Model Output | |
| │ | |
| Citations | |
| ``` | |
| Potential checks include: | |
| - source provenance, | |
| - access permissions, | |
| - retrieval relevance, | |
| - outdated documents, | |
| - unsupported claims, | |
| - source-to-answer traceability. | |
| --- | |
| # Auditing open and open-weight models | |
| Open models create additional opportunities for technical inspection. | |
| An audit may examine: | |
| - model provenance, | |
| - repository history, | |
| - model files, | |
| - license, | |
| - architecture, | |
| - dependency chain, | |
| - quantized variants, | |
| - adapters, | |
| - fine-tunes, | |
| - evaluation artifacts, | |
| - deployment configuration. | |
| Open access does not automatically mean that a system is trustworthy, compliant or safe. | |
| It does, however, enable forms of inspection that may not be possible with a fully closed model. | |
| --- | |
| # AI auditing and security | |
| AI auditing and AI security overlap but serve different functions. | |
| Security asks: | |
| > **How do we protect the system?** | |
| Auditing asks: | |
| > **Can we verify whether the expected protections existed and worked?** | |
| Security evidence may include: | |
| - access logs, | |
| - permission changes, | |
| - secret rotation, | |
| - authentication events, | |
| - sandbox violations, | |
| - network access, | |
| - anomalous behavior, | |
| - incident records. | |
| Auditing can use this evidence to test whether controls were operating as intended. | |
| --- | |
| # AI auditing and governance | |
| Governance defines responsibilities, rules and decision structures. | |
| Auditing examines whether those structures are implemented and evidenced. | |
| A governance policy might say: | |
| > High-impact agent actions require human approval. | |
| An audit might test: | |
| 1. Which actions are classified as high impact? | |
| 2. Was the approval mechanism technically enforced? | |
| 3. Are approval events logged? | |
| 4. Can approvals be linked to the corresponding action? | |
| 5. Were any actions executed without approval? | |
| This turns governance language into testable controls. | |
| --- | |
| # Framework alignment | |
| AI auditing does not need to depend on one framework. | |
| Audit programs may map evidence and controls to internal requirements or external frameworks. | |
| Examples include: | |
| - **NIST AI Risk Management Framework (AI RMF)** | |
| - **ISO/IEC 42001 AI Management Systems** | |
| - organizational policies, | |
| - sector-specific requirements, | |
| - contractual obligations, | |
| - applicable laws and regulations. | |
| NIST describes its AI RMF as a voluntary framework for managing AI risk throughout the lifecycle, while ISO/IEC 42001 specifies requirements for establishing and continually improving an AI management system. | |
| Auditing can provide evidence for checking whether selected requirements and controls have actually been implemented. | |
| --- | |
| # Audit metrics | |
| No single metric determines whether an AI system is auditable. | |
| Useful operational indicators may include: | |
| | Metric | What it can indicate | | |
| |---|---| | |
| | Trace Coverage | Share of relevant executions with complete traces | | |
| | Evidence Completeness | Availability of required audit fields | | |
| | Model Version Coverage | Ability to identify model versions | | |
| | Tool Call Traceability | Ability to reconstruct tool executions | | |
| | Approval Coverage | Required actions with recorded approval | | |
| | Unexplained Action Rate | Actions lacking sufficient evidence | | |
| | Control Failure Rate | Frequency of failed controls | | |
| | Override Rate | Frequency of human overrides | | |
| | Audit Finding Closure | Remediation progress | | |
| | Time to Reconstruct | Effort required to reconstruct an event | | |
| These are examples, not universal standards. | |
| --- | |
| # Practical audit questions | |
| ## Model | |
| - Which model is running? | |
| - Which version? | |
| - Who approved it? | |
| - What evaluation evidence exists? | |
| - Has it changed since the previous review? | |
| ## Agent | |
| - What goal was assigned? | |
| - Which decisions were autonomous? | |
| - Which tools were available? | |
| - Which permissions existed? | |
| - Which actions required approval? | |
| ## Data | |
| - Which sources were accessed? | |
| - Was the agent authorized to access them? | |
| - Can retrieved information be traced? | |
| - Were sensitive data controls applied? | |
| ## Runtime | |
| - Are executions logged? | |
| - Are traces complete? | |
| - Can events be correlated? | |
| - Are important configuration changes recorded? | |
| ## Controls | |
| - Which controls should have activated? | |
| - Did they activate? | |
| - Can this be demonstrated with evidence? | |
| ## Outcome | |
| - Did the system achieve its goal? | |
| - Did it violate any constraints? | |
| - Was human correction required? | |
| - Were unexpected side effects created? | |
| --- | |
| # Planned Hugging Face resources | |
| The **Auditing** organization is intended to develop practical resources around AI assurance and auditability. | |
| Potential projects include: | |
| ## AI Audit Explorer | |
| An interactive guide to: | |
| - audit scopes, | |
| - evidence types, | |
| - controls, | |
| - risks, | |
| - audit procedures, | |
| - findings. | |
| ## Agent Audit Trail Explorer | |
| Visualize the execution path of an AI agent: | |
| ```text | |
| Prompt | |
| → Plan | |
| → Tool | |
| → Permission | |
| → Action | |
| → Observation | |
| → Decision | |
| → Result | |
| ``` | |
| ## AI Audit Checklist | |
| Generate structured audit questions based on: | |
| - models, | |
| - agents, | |
| - RAG, | |
| - tool use, | |
| - sensitive data, | |
| - deployment architecture. | |
| ## Audit Evidence Schema | |
| A reusable data schema for storing evidence such as: | |
| - model versions, | |
| - tool calls, | |
| - approvals, | |
| - traces, | |
| - policy checks, | |
| - findings. | |
| ## AI Audit Report Builder | |
| Turn structured evidence into a consistent technical audit report. | |
| ## Agent Control Matrix | |
| Map agent capabilities to: | |
| - risks, | |
| - permissions, | |
| - controls, | |
| - evidence, | |
| - testing procedures. | |
| --- | |
| # Example audit record | |
| ```json | |
| { | |
| "audit_event_id": "evt_001", | |
| "timestamp": "2026-09-25T08:30:00Z", | |
| "agent_id": "procurement-agent", | |
| "session_id": "session_8842", | |
| "model": "model-name", | |
| "model_version": "version-id", | |
| "goal": "Compare approved suppliers", | |
| "tool": "supplier-database", | |
| "action": "read", | |
| "permission": "supplier.read", | |
| "human_approval_required": false, | |
| "policy_checks": [ | |
| "approved-source-only" | |
| ], | |
| "result": "success", | |
| "trace_id": "trace_3902" | |
| } | |
| ``` | |
| This is only an illustrative schema. | |
| Real systems may require different fields, stronger integrity protections and additional privacy controls. | |
| --- | |
| # Audit finding structure | |
| A useful finding should be more specific than: | |
| > “The AI system needs better monitoring.” | |
| A structured finding might contain: | |
| ```text | |
| Finding ID | |
| Title | |
| System / Component | |
| Observed Condition | |
| Expected Control | |
| Evidence | |
| Risk / Impact | |
| Root Cause | |
| Recommendation | |
| Owner | |
| Due Date | |
| Status | |
| Re-test Result | |
| ``` | |
| This makes findings actionable and reviewable. | |
| --- | |
| # Auditability maturity | |
| Organizations can think about auditability as a progression. | |
| ```text | |
| LEVEL 1 — UNKNOWN | |
| Little structured evidence | |
| LEVEL 2 — LOGGED | |
| Basic events are recorded | |
| LEVEL 3 — TRACEABLE | |
| Important decisions and actions can be reconstructed | |
| LEVEL 4 — CONTROLLED | |
| Controls are mapped to verifiable evidence | |
| LEVEL 5 — CONTINUOUS | |
| Selected controls and evidence are continuously assessed | |
| ``` | |
| This is a practical project model, not an official maturity standard. | |
| --- | |
| # Principles | |
| ## Evidence before assumptions | |
| Audit conclusions should be based on available evidence. | |
| ## Traceability before opacity | |
| Important actions should be reconstructable. | |
| ## Version everything important | |
| Models, prompts, tools, policies and configurations change. | |
| ## Audit actions, not only outputs | |
| Agent systems should be examined at the trajectory level. | |
| ## Separate monitoring from assurance | |
| Monitoring generates signals. Auditing tests evidence and controls. | |
| ## Risk-based scope | |
| Not every AI system requires the same level of audit depth. | |
| ## Independent verification | |
| Where appropriate, evidence should be reviewable independently of the component being audited. | |
| ## Reproducibility | |
| Important audit tests should be repeatable where technically possible. | |
| --- | |
| # Who this organization is for | |
| Auditing is intended for: | |
| - AI engineers, | |
| - ML engineers, | |
| - agent developers, | |
| - AI platform teams, | |
| - security teams, | |
| - governance teams, | |
| - risk professionals, | |
| - internal audit teams, | |
| - compliance teams, | |
| - MLOps teams, | |
| - enterprise architects, | |
| - researchers, | |
| - organizations deploying autonomous AI systems. | |
| --- | |
| # Why Hugging Face? | |
| AI auditing increasingly requires technical artifacts rather than policy documents alone. | |
| Hugging Face provides a natural environment for combining: | |
| ```text | |
| Models | |
| + | |
| Datasets | |
| + | |
| Spaces | |
| + | |
| Collections | |
| + | |
| Evaluation artifacts | |
| + | |
| Technical documentation | |
| ``` | |
| This makes it possible to create practical and reproducible resources for AI auditing. | |
| The goal of **Auditing** is to help connect technical AI systems with **evidence, control and assurance**. | |
| --- | |
| # Reference frameworks | |
| The project may use established frameworks as reference points while remaining framework-independent. | |
| - NIST AI Risk Management Framework | |
| https://www.nist.gov/itl/ai-risk-management-framework | |
| - NIST AI Resource Center | |
| https://airc.nist.gov/ | |
| - ISO/IEC 42001 — Artificial intelligence management systems | |
| https://www.iso.org/standard/42001 | |
| These references do not replace organization-specific, sector-specific or legal requirements. | |
| --- | |
| # Collaboration | |
| We welcome collaboration around: | |
| - AI auditing, | |
| - agent auditability, | |
| - AI assurance, | |
| - audit trails, | |
| - traceability, | |
| - model validation, | |
| - observability, | |
| - governance, | |
| - technical controls, | |
| - evaluation, | |
| - AI security, | |
| - audit datasets, | |
| - open audit schemas, | |
| - Hugging Face Spaces, | |
| - enterprise AI governance tooling. | |
| **Collaboration:** agenten@magenta.de | |
| --- | |
| # Independent project | |
| **Auditing** is an independent technical project. | |
| It is not an official Hugging Face organization and is not affiliated with NIST, ISO or any regulator, certification body, audit firm or model provider unless explicitly stated. | |
| The project does not provide legal, regulatory, certification or professional audit advice. | |
| Requirements differ by jurisdiction, sector and use case. Important compliance or certification decisions should be verified with the relevant qualified professionals and authoritative sources. | |
| --- | |
| # Auditing | |
| ### **Evidence for AI. Traceability for agents. Assurance for autonomous systems.** | |
| AI is moving from systems that generate outputs to systems that can **decide, delegate and act**. | |
| As autonomy grows, the ability to reconstruct and verify those actions becomes increasingly important. | |
| **Auditing exists to make AI systems more examinable, traceable and accountable — from models to agents to the infrastructure around them.** | |