--- tags: - autoencoder - formal-logic - experimental - ipfs-datasets-py --- # security-ir: experimental 384D checkpoint **Held-out exact IR reconstruction: 0/2. Native structural validity: 2/2.** These two authored paraphrase examples share semantic targets with training; they are not independent real-world benchmark tasks. The model generates plausible but semantically incorrect candidates and must not authorize actions. This is an opt-in development artifact. It is **not the default symbolic compiler**, not a verified formalizer, and not a security repair policy. All qualification, admission and proof-authority flags remain false. No LLM/provider fallback runs. ## Runtime Input embeddings use `thenlper/gte-small` at commit `17e1f347d17fe144873b1201da91788898c639cd` (384 dimensions). The embedding encoder weights are not included. Source inference requires the corresponding verified local encoder assets and rejects inputs above 512 tokens without truncation. The installed `ipfs_datasets_py` implementation must match the source hashes in the checkpoint; incompatible sources are rejected. This development runtime is currently implemented in the matching workspace checkout, not claimed to be in an existing PyPI release. No repository-supplied Python is executed on download. ```python from ipfs_datasets_py.logic.security_ir import open_autoencoder model = open_autoencoder() # library catalog pins a complete Hub commit + SHA256 report = model.infer([{"id": "sample", "source_text": source, "embedding": vector384}]) # Or: model.infer_texts([source], snapshot_path=verified_gte_snapshot) ``` ## Training and scope Security, Intent and UI/UX heads inherit all compatible nonlexical tensors from a trained local Legal 384D head. New lexical rows use trained parent row means; no random parameters remain in their initial transferred state. Each native head used 12 authored training rows, 2 tuning rows, 500 epochs / 1000 optimizer updates; selection used tuning loss only. Outputs are explicitly local typed fragments, not complete verified programs or native documents. Security operator fragments still require their enclosing program and source-reference context. Legal retains its separate sparse core and learned formula head, with parser-derived features. These checkpoints were **not trained on the complete CVEFixes or SkillCenter corpora**. The separate full-feature structural-family experiment (15/9/8/6 routes) does not measure this 384D source decoder and is not its training loss or coverage. Weight/conditioning ablations demonstrate numerical dependence, not correct meaning. ## Provenance and license The manifest records the actual parent hashes, training manifests, embedding assets and measured validation. Source screening found no private documents, credentials, benchmark solutions or third-party corpus text in the new authored fixture export. Original Legal files and the upstream Hugging Face datasets were not overwritten. This release makes no new license grant for inherited weights. The runtime source repository retains its AGPL-3.0 license; the new authored fixture export separately declares Apache-2.0. A corpus license is not asserted to be the weight license. The checkpoint and evidence are immutable under `releases/20261002-384-development-v1/`. Manifest SHA256: `d9fc241c7c6a039068681fab18b6a36fbcbd1752c2683ba1dc16423ce94aba2a`. ## Additional verified development releases (2026-10-01) The following opt-in structured profiles use a fixed typed JSON schema and learned scalar vocabulary. Their cards distinguish raw reconstruction, hybrid AST normalization, and compilation. They do not replace the symbolic formalizers or confer proof authority. - [structured_native_v3](https://huggingface.co/Publicus/security-ir-autoencoder/tree/e451ce831c4fbc9079b7516e5d415dbab9e219ae/releases/20261001-structured-native-v3-experimental) - [security_styles_v4_diagnostic](https://huggingface.co/Publicus/security-ir-autoencoder/tree/e451ce831c4fbc9079b7516e5d415dbab9e219ae/releases/20261001-security-styles-v4-diagnostic) [Runtime code and exact release descriptors](https://github.com/endomorphosis/ipfs_datasets_py/blob/f4e28be225455add0f29259fa6ecf238d23c6316/docs/autoencoders/experimental_structured_checkpoints.md). All new package hashes and cold/offline numerical replay were verified.