File size: 8,439 Bytes
e91f9b4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
cfffa10
e91f9b4
cfffa10
 
e91f9b4
cfffa10
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
e91f9b4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
# F-17 Profiling β€” Parallelizing the 30 Signals

## Method

27 of the 30 signals (19-signal statistical bundle + 8 classical forensic detectors +
`image_type_classifier`) are pure numpy/opencv/scipy/scikit-image β€” no torch dependency β€” so they
were timed for real: a synthetic 1600Γ—1200 JPEG (smooth gradient + sensor-like noise, quality 88),
5 runs each, mean/min/max recorded. The 3 torch-dependent signals (DIRE, CLIP, own-embedding)
couldn't be executed in the profiling session's sandbox (no working torch install there) β€” their
relative cost was estimated by reading the model architecture directly instead.

## Real measured numbers (mean of 5 runs, synthetic 1600Γ—1200 JPEG)

| Signal | Mean | Notes |
|---|---|---|
| Statistical bundle (19 signals, 1 call) | 2116.5ms | Internally: BasicSignals 1396.7ms, UltraSignals 264.5ms, AdvancedSignals 232.2ms, CovarianceSignals 149.8ms |
| JPEG ghost | 1026.1ms | |
| DCT frequency | 758.6ms | |
| Noiseprint | 557.7ms | |
| Noise map | 481.6ms | |
| CFA | 154.8ms | |
| Image type classifier | 126.4ms | Gates prnu/ela/metadata β€” must resolve before those three, independent of everything else |
| PRNU | 92.0ms | |
| ELA | 75.2ms | |
| Metadata | 0.1ms | |

**Sum if sequential: ~5.4s. Slowest single signal (the parallel floor for this group): ~2.1s
(statistical bundle).**

Within the statistical bundle itself, F-16 already decomposed it into 4 independent components
sharing only a read-only `ImageContext` β€” `BasicSignals` alone is ~66% of the bundle's own total
(1.4s of 2.1s), so there's a second, smaller layer of parallelization available inside the bundle
if the outer-level win alone isn't enough later. Not implemented in this pass β€” flagged for
whoever picks up further F-17 work.

## DIRE / CLIP / own-embedding β€” structural estimate, not measured

Read directly from `dire_detector.py`, `clip_detector.py`, `own_embedding_detector.py`:

- **DIRE**: loads a full Stable Diffusion 2.1 pipeline, runs **20 sequential UNet forward passes**
  (~860M params) at 512Γ—512, plus a VAE encode/decode. Almost certainly the dominant cost of the
  whole pipeline on CPU β€” likely an order of magnitude past the 2.1s statistical bundle.
- **CLIP**: one forward pass, ViT-B/32 (~88M params). Fast relative to DIRE.
- **own-embedding**: one forward pass, EfficientNet-B0 (~5M params). Fastest of the three.

To get a real DIRE number, run this in an environment with working torch (this repo's own dev
environment, not necessarily the profiling sandbox):

```python
import time, statistics
from backend.services.dire_detector import DIREDetector

with open("some_real_test_image.jpg", "rb") as f:
    image_bytes = f.read()

times = []
d = DIREDetector()
d.detect(image_bytes, "warmup.jpg")  # first call also pays model-load cost β€” exclude it
for _ in range(3):
    t0 = time.perf_counter()
    d.detect(image_bytes, "test.jpg")
    times.append(time.perf_counter() - t0)
print(f"DIRE mean: {statistics.mean(times):.2f}s  (min {min(times):.2f}s, max {max(times):.2f}s)")
```

## The scheduler-sharing risk β€” RESOLVED

DIRE's `DDIMScheduler` was loaded once and cached process-wide (`ModelCache`, F-7) β€”
`cached_model['scheduler']` β€” with every `DIREDetector` instance sharing that *same object*. Both
`_add_noise()` and the denoising loop call mutating methods on it (`scheduler.add_noise()`,
`scheduler.set_timesteps()`, `scheduler.step()`). `denoise_steps` is hardcoded to `20`, so a
`set_timesteps(20)` race between two concurrent DIRE calls likely produced the same values either
way β€” probably not silently wrong in practice. But that was incidental, not a guarantee, and it was
exactly the risk the original F-17 planning flagged as "a plausible interaction risk nobody has
tested yet."

**Fixed** (follow-up commit, after this document was first written): `DIREDetector._load_model()`
now clones a fresh scheduler from the cached one's config on every cache hit
(`type(cached_scheduler).from_config(cached_scheduler.config)`) instead of sharing the cached
object directly. `from_config()` does no disk/network I/O β€” just Python object construction from a
config dict β€” so this is negligible overhead per request. `self.pipe` (the actual UNet/VAE weights)
stays shared: frozen, eval-mode forward passes don't mutate shared state, so sharing that part
remains safe. Verified directly: two `DIREDetector` instances that both hit the cache now get two
distinct scheduler objects (`first.scheduler is not second.scheduler`); confirmed this is a real
fix, not just plausible, by reverting it and watching the new regression test fail with exactly the
shared-object assertion error, then pass again once restored.

**This does not by itself mean DIRE should join the parallel signal pool.** The remaining open
question is the one this document's "DIRE / CLIP / own-embedding β€” structural estimate, not
measured" section above already raised: there's still no real DIRE timing number, so it's not known
whether DIRE dominates total latency so heavily that parallelizing it alongside anything else
provides only a marginal win, or whether there's a genuine, substantial gain available. Get a real
number (the script above) before deciding whether extending the thread pool to DIRE/CLIP/
own-embedding is worth the added complexity β€” the thread-safety blocker is gone, but the
latency-value question isn't answered yet.

## What was actually shipped this round (partial F-17)

The 9 mutually-independent, non-torch signal calls β€” the statistical bundle + 8 classical
forensic detectors β€” now run in a `ThreadPoolExecutor` inside `AdvancedEnsembleDetector.detect()`,
instead of one after another. Every module involved was grepped for module-level caches/globals
before this change; none exist (all are pure functions over their own local data), so there's no
DIRE-scheduler-style shared-state risk in this half of the change.

**A real trap found and fixed during verification**: the profiling/development sandbox turned out
to have exactly 1 CPU core. On 1 core, 9 threads contending for that core measured **~1.5x SLOWER**
than plain sequential β€” pure context-switch overhead, zero real parallelism, not a hypothetical
concern. Fixed by capping `max_workers = min(len(_signal_tasks), max(1, os.cpu_count() or 1))`, so
the pool degrades to one-thread-at-a-time (matching sequential performance) on a constrained host
instead of regressing, and scales up to real concurrency on whatever's actually available.

The project's actual production target (Hugging Face Spaces CPU Basic) is confirmed 2 vCPU β€” real,
though bounded (not the full 9-way ideal), benefit is expected there. DIRE/CLIP/own-embedding
remain sequential, tracked as a separate follow-up gated on the scheduler-sharing question above
and a real DIRE number from the script in this doc.

## Verification performed

- Ran the real, patched `AdvancedEnsembleDetector.detect()` (torch import stubbed just enough to
  satisfy `dire_detector.py`/`clip_detector.py`'s type annotations at import time; `dire_detector`/
  `clip_detector`/`own_detector` swapped for instant fakes so timing measures only the 9-signal
  pool this change touches) β€” 3 repeated runs produced identical `ai_probability` and identical
  30-signal sets, confirming no race condition/data corruption from the thread pool.
- Directly verified the `max_workers` capping logic against 3 scenarios: `cpu_count()=1` β†’ 1
  worker, `cpu_count()=64` β†’ 9 workers (capped at task count, not core count), `cpu_count()=None`
  β†’ 1 worker (not passed through as `None`, which `ThreadPoolExecutor` would otherwise interpret
  as "pick a default based on core count," silently reintroducing the oversubscription bug this
  fix exists to prevent).
- Full `backend/` still `py_compile`-clean after the change.
- Could **not** measure real multi-core speedup in the profiling sandbox (1 core, by definition
  can't demonstrate parallelism). Recommend running the timing harness below on this repo's own
  environment (the 2-vCPU HF Space, or local dev hardware) to confirm the real-world win:

```python
import time
from backend.services.advanced_ensemble_detector import AdvancedEnsembleDetector

with open("some_real_test_image.jpg", "rb") as f:
    image_bytes = f.read()

t0 = time.perf_counter()
AdvancedEnsembleDetector(image_bytes, "test.jpg").detect()
print(f"Full detect(): {time.perf_counter() - t0:.2f}s")
```