File size: 5,007 Bytes
c00473f
 
 
 
 
 
6474d5b
 
 
c00473f
 
 
 
 
 
 
 
 
 
 
6474d5b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c00473f
 
 
 
 
a30407f
 
c00473f
 
 
 
 
 
 
 
 
 
 
6474d5b
 
 
c00473f
 
 
6474d5b
c00473f
 
6474d5b
c00473f
 
 
 
6474d5b
 
 
c00473f
 
 
 
 
6474d5b
 
 
 
c00473f
 
6474d5b
 
 
 
 
 
1ad9dd1
6474d5b
1ad9dd1
6474d5b
 
 
 
1ad9dd1
6474d5b
1ad9dd1
 
 
 
 
 
 
 
 
 
 
 
6474d5b
 
 
 
 
 
 
c00473f
 
 
6474d5b
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
---
license: mit
language:
- en
base_model: microsoft/deberta-v3-base
pipeline_tag: text-classification
metrics:
- f1
- accuracy
tags:
- text-classification
- deberta-v3
- spatial-language
- spatial-reasoning
library_name: transformers
---

# DeBERTa-v3-base — Spatial Language Detection

A fine-tuned [`microsoft/deberta-v3-base`](https://huggingface.co/microsoft/deberta-v3-base)
for **word-level spatial-language detection**: given an utterance and a **target word**, it
decides whether that word is being used as spatial language (location, direction, or a spatial
relationship) **in that context**`1` = spatial, `0` = not. The same word can be spatial in
one utterance ("go *up* the ramp") and not in another ("what's *up*?"), so the model always
judges a word together with its sentence.

For the full pipeline (dictionary gating, calibrated confidence, evaluation) and example
datasets, see the GitHub repo:
https://github.com/SamAgnoli/spatial-language-classifier

## Input format

This is a **sentence-pair** classifier: pass the **utterance** as the first segment and the
**target word** as the second — `tokenizer(utterance, target_word)`. Passing a whole sentence
on its own is *not* how the model was trained and gives unreliable results.

## Labels

| id | label | meaning |
|----|-------------|------------------------------------------|
| 0  | not_spatial | word is not spatial language |
| 1  | spatial     | word is spatial language |

## Usage

```python
from transformers import AutoModelForSequenceClassification, AutoTokenizer
import torch

model_id = "SamAgnoli/deberta-v3-base-spatial-language-detection"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)

utterance   = "The cat is sitting on top of the bookshelf."
target_word = "top"                       # the word you want judged
inputs = tokenizer(utterance, target_word, return_tensors="pt", truncation=True)
with torch.no_grad():
    logits = model(**inputs).logits
pred = logits.argmax(-1).item()
print(model.config.id2label[pred])        # -> "spatial"
```

Or with a pipeline (note the `text` / `text_pair` keys):

```python
from transformers import pipeline

clf = pipeline("text-classification",
               model="SamAgnoli/deberta-v3-base-spatial-language-detection")
print(clf({"text": "The cat is sitting on top of the bookshelf.", "text_pair": "top"}))
```

## Training

- **Base model:** `microsoft/deberta-v3-base`
- **Task:** binary sentence-pair classification (a word, in its utterance, spatial vs. not)
- **Split:** group-aware 70/15/15 by speaker session (no session spans splits)
- **Hyperparameters:** 2 epochs · lr 2e-5 · batch 16 · weight decay 0.01 · warmup 0.1 ·
  max_length 128 · fp16 · seed 42 · best checkpoint by F1
- **Framework:** 🤗 Transformers

## Evaluation (held-out test set)

Reported for two views: **dictionary candidates only** (the meaningful view — words a spatial
dictionary flags as plausibly spatial) and **overall** (every word, dominated by trivially
non-spatial tokens).

### Candidates only — 799 words (570 not-spatial, 229 spatial)

| class | precision | recall | F1 | support |
|---|---|---|---|---|
| not_spatial | 0.958 | 0.925 | 0.941 | 570 |
| spatial | 0.827 | 0.900 | 0.862 | 229 |

**Overall accuracy: 0.917** (733 of 799 words correct) · macro-F1 0.901 · Cohen's κ 0.803

Reading it per class: the model catches **90.0%** of truly-spatial words (recall) at **82.7%**
precision; for non-spatial words it's 92.5% recall at 95.8% precision. (In clinical terms:
sensitivity 0.900, specificity 0.925, PPV 0.827, NPV 0.958.)

### Overall — every word, 5,207 tokens

accuracy **0.987** · spatial-F1 **0.862** · Cohen's κ **0.855**

> The two views share the *same* spatial predictions (229 spatial words, same 206 caught). Only
> the non-spatial pool differs, which is why "overall" accuracy looks higher — it's padded with
> ~4,400 easy non-candidate words the model trivially gets right. Judge the model by the
> **candidates-only** view.

## Calibration

Raw probabilities are over-confident, so a post-hoc **temperature scaling** factor
(**T = 1.816**, fit on the validation candidates) rescales them into a calibrated `P(spatial)`
you can read literally. Temperature scaling is monotonic, so the hard 0/1 decision is unchanged.
See section 7.5 of the repo.

## Intended use & limitations

- Built for **per-word** spatial judgments within an English utterance. In production it is
  paired with a spatial-dictionary gate that selects candidate words; words the dictionary
  misses (e.g., misspellings) are never sent to the model.
- Trained on parent–child tinkering-reflection speech — performance on other domains, genres,
  or languages is not guaranteed.
- The data is strongly imbalanced (~4% spatial overall); judge quality by the
  **candidates-only** view, not the overall numbers.
- Review predictions before relying on them in downstream systems.