snapkitty
machine-learning
python
File size: 5,080 Bytes
2af8b7a
 
 
 
 
 
28ded58
 
2af8b7a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
---
license: other
license_name: snapkitty-tri-license
license_link: https://huggingface.co/Snapkitty/pocketlearn/blob/main/LICENSE.tri
tags:
- snapkitty
- machine-learning
- python
---

> Source: [github.com/SNAPKITTYWEST/pocketlearn](https://github.com/SNAPKITTYWEST/pocketlearn)

# PocketLearn

**Symbolic cognitive architecture: XML + XSLT + ILP + ASP + FORTH. Zero Python.**

Learn = build a visible theory.

Neural net: `learn = adjust W -= lr * grad`. Knowledge disappears into numbers you can't read.

This: `learn = build a visible theory.`

---

## What it does

```
sample_corpus.txt
      |
      v
corpus_tokens.xml          (tokenizer β€” 69 tokens, 52 vocab)
      |
      v
ontology.xml               (seed concepts: stack_op, compiler_word, meta_word...)
      |
      +--[XSLT]----------> background.pl          (Prolog co-occurrence facts)
      |
      +--[XSLT]----------> ontology_induction_generated.pl   (ILP engine, GENERATED by XSLT)
                                    |
                                    v
                             swipl learns rules:
                             Induced: is_a(W, stack_op) :- cooccur(W, 'drop'). F1=0.60
                             Propose: include should be is_a(stack_op) cnt=1
                                    |
                                    v
                           ontology_induced.xml    (updated ontology with induced members)
                                    |
                     +------[XSLT]-+------[XSLT]--+
                     |                            |
                     v                            v
             ASP validation               generated_corpus_induced.fth
             clingo rejects               gforth runs the learned dictionary
             contradictions
             (dup = stack_op AND
              compiler_word -> UNSAT)
```

**The meta-trick:** `ontology_to_induction.xslt` generates the Prolog ILP engine from `ontology.xml`. So the whole system is self-describing β€” XSLT generates Prolog that learns rules from XML co-occurrence stats.

---

## Run

```bash
# Install (Mac)
brew install libxslt swi-prolog clingo gforth

# Install (Linux)
sudo apt install -y xsltproc swi-prolog gringo gforth

# Build β€” full pipeline
make

# Run the FORTH (pre-built, no deps needed)
make demo-prebuilt
```

---

## What you get

```bash
make
# [3/7] ILP engine via XSLT
# [4/7] ILP Induction
# Induced: is_a(W, stack_op)       :- cooccur(W, 'drop').       F1=0.60
# Induced: is_a(W, compiler_word)  :- cooccur(W, 'semicolon').  F1=0.75
# Induced: is_a(W, learning_word)  :- cooccur(W, 'statistical'). F1=0.80
# Proposing: include  should be is_a(stack_op)      (cooccurs with 'drop')
# Proposing: defined  should be is_a(compiler_word) (cooccurs with 'semicolon')
# Proposing: similarity should be is_a(learning_word)
# [5/7] ASP: SATISFIABLE
# [6/7] FORTH written

make demo
# PocketLearn FORTH β€” seed + ILP-induced vocab
# vocab size: 18
# Induced: include (by drop), defined (by semicolon), similarity (by statistical)
```

---

## Files

| File | Role |
|------|------|
| `sample_corpus.txt` | Input text |
| `corpus_tokens.xml` | Tokenized corpus (XML) |
| `ontology.xml` | Seed concepts with members + co-occurrence strengths |
| `ontology_induced.xml` | Output ontology with ILP-induced members |
| `corpus_to_background.xslt` | XML β†’ Prolog co-occurrence facts |
| `ontology_to_induction.xslt` | **Generates** the Prolog ILP engine from ontology.xml |
| `ontology_to_asp.xslt` | XML β†’ ASP validation facts |
| `corpus_to_forth.xslt` | XML β†’ FORTH dictionary |
| `ontology_induction_generated.pl` | ILP engine (XSLT output) β€” run with swipl |
| `generated_corpus_induced.fth` | Final FORTH (seed + induced) β€” run with gforth |
| `ontology.asp` | ASP contradiction rules |
| `Makefile` | Full pipeline |

---

## Why this instead of a transformer

| | Transformer | PocketLearn |
|--|--|--|
| Inspectable | No β€” weights are numbers | Yes β€” open `ontology_induced.xml` |
| Reproducible | No β€” depends on random seed | Yes β€” same XML = same FORTH, bit-for-bit |
| Debuggable | No | Yes β€” stack blow β†’ trace to corpus_tokens.xml line β†’ XSLT template |
| Hallucinates | Yes β€” `dup = delete` possible | No β€” ASP kills contradictions |
| Learns deep semantics | Yes | No |

It won't discover deep semantics. It will never hallucinate `dup = delete` because ASP kills it.

---

**Ahmad Ali Parr Β· Bel Esprit D'Accord Irrevocable Trust Β· EIN 42-697643**

`Omega = TRUST AND CODE`

---

## License

Licensed under **SnapKitty Tri-License**. Full text: [LICENSE.tri](LICENSE.tri).

### πŸ’Ό Commercial License

Snapkitty code is free and open under **AGPL-3.0** for open-source use. Building a commercial product or service? A **proprietary commercial license** from Snapkitty Collective LLC lets you ship this code without the AGPL's source-sharing and network-use obligations.

**[β†’ Get a commercial license](mailto:A.parr@belespritdaccord.uk?subject=Commercial%20license:%20pocketlearn)** Β· A.parr@belespritdaccord.uk