Spaces:
Sleeping
Sleeping
File size: 24,569 Bytes
1605cbb | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 | # FALSIFY — Requirements & Build Contract
> **Tagline:** Drop one contradicting fact; watch dependent conclusions die and a losing hypothesis rise — permanently, across sessions.
FALSIFY is a belief-revision research copilot. It treats a research inquiry as a **living belief graph** (not a chat log). When new evidence contradicts an existing evidence node, FALSIFY performs **belief revision**: it flips the contradicted evidence to `refuted`, propagates the refutation **forward** through the dependency chain, marks every dependent conclusion invalid, promotes the best surviving hypothesis, and **surgically forgets** the orphaned dead-ends. The refutation is written to graph node truth-state so the **next session's `recall()` skips killed branches** — while a plain-RAG baseline still cites the stale fact.
This document is the **contract** for all downstream build agents. Every API named here has been verified against the read-only reference at `/workspaces/cognee`. Line references are cited where load-bearing.
---
## 0. Verified Cognee API Surface (grep-confirmed)
| API | Exact import / call | Location (reference, read-only) |
|---|---|---|
| `Task` | `from cognee.modules.pipelines.tasks.task import Task` **(lowercase module `task`)** | `cognee/modules/pipelines/tasks/task.py` |
| `memify` | `import cognee; await cognee.memify(extraction_tasks=[...], enrichment_tasks=[...], data=..., node_type=NodeSet, node_name=[...])` | `cognee/modules/memify/memify.py:26` |
| `remember` | `await cognee.remember(data, dataset_name=..., session_id=...)` | `cognee/api/v1/remember/remember.py:623` |
| `recall` | `await cognee.recall(query_text, query_type=..., session_id=..., top_k=...)` | `cognee/api/v1/recall/recall.py:361` |
| `improve` | `await cognee.improve(...)` | `cognee/api/v1/improve/improve.py:36` |
| `forget` | `await cognee.forget(...)` | `cognee/api/v1/forget/forget.py:16` |
| `get_graph_engine` | `from cognee.infrastructure.databases.graph import get_graph_engine` → `ge = await get_graph_engine()` **(async)** | `cognee/infrastructure/databases/graph/get_graph_engine.py:110` |
| `get_vector_engine` | `from cognee.infrastructure.databases.vector import get_vector_engine` → `ve = get_vector_engine()` **(sync)** | `cognee/infrastructure/databases/vector/get_vector_engine.py:44` |
| `DataPoint`, `Embeddable`, `Dedup`, `LLMContext` | `from cognee.infrastructure.engine import DataPoint, Embeddable, LLMContext, Dedup` | `cognee/infrastructure/engine/models/DataPoint.py:27` |
| `LLMGateway` | `from cognee.infrastructure.llm.LLMGateway import LLMGateway` | `cognee/infrastructure/llm/LLMGateway.py:52` |
### Graph engine methods (verified in `graph_db_interface.py` + `ladybug/adapter.py`)
```python
ge = await get_graph_engine()
# Forward traversal — THE propagation rail. Returns (nodes, edges).
nodes, edges = await ge.get_neighborhood(node_ids: List[str], depth: int = 1, edge_types: Optional[List[str]] = None)
# graph_db_interface.py:303 ; ladybug/adapter.py:2068
# Immediate neighbors of ONE node. Returns List[Tuple[source_props, edge_props, target_props]].
conns = await ge.get_connections(node_id: Union[str, UUID])
# graph_db_interface.py:289 ; ladybug/adapter.py:1859
# Truth-state: the alive/dead flag recall() filters on. PERSISTS across restart.
await ge.set_node_truth_state({node_id: {"truth_alignment": ["refuted"], "truth_epoch": N}})
state = await ge.get_node_truth_state(node_ids: List[str])
# graph_db_interface.py:357,364 ; ladybug impl 1687,1712 ; writes props "truth_alignment" (list) + "truth_epoch" (int) — adapter.py:1576-1578
# Confidence / health signal.
await ge.set_node_feedback_weights({node_id: 0.9}) # graph_db_interface.py:348
w = await ge.get_node_feedback_weights(node_ids) # graph_db_interface.py:341
# Edges.
await ge.add_edge(from_node: str, to_node: str, relationship_name: str, edge_properties: Dict = {}) # ladybug/adapter.py:1222
# Surgical delete.
await ge.delete_nodes(node_ids: List[str]) # graph_db_interface.py:102 ; ladybug 1042
```
### Vector engine methods (verified in `LanceDBAdapter.py`)
```python
ve = get_vector_engine()
hits = await ve.search(collection_name: str, query_text: str = None, limit: int = 15, include_payload: bool = False, ...)
# LanceDBAdapter.py:992 — each hit exposes `.id`, `.score` (COSINE DISTANCE, lower = closer), `.payload`
await ve.delete_data_points(collection_name: str, data_point_ids: List[UUID])
# vector_db_interface.py:160 ; LanceDBAdapter.py:1106
```
### LLMGateway (verified `LLMGateway.py:59`)
```python
# STATIC method — do NOT instantiate. Returns a coroutine; await it.
result = await LLMGateway.acreate_structured_output(
text_input="A(...): <claim>\nB(...): <claim>",
system_prompt="skeptical analyst ...",
response_model=ContradictionJudgement, # a pydantic BaseModel
)
```
### DataPoint subclass pattern (verified `Tool.py`)
```python
from typing import Annotated
from cognee.infrastructure.engine import DataPoint, Embeddable, Dedup
class Hypothesis(DataPoint):
statement: Annotated[str, Embeddable(), Dedup()]
question_id: str
# identity_fields in metadata dedups nodes across sessions (see §2.4)
metadata: dict = {"index_fields": ["statement"], "identity_fields": ["question_id", "statement"]}
```
> **NOTE for implementers:** The FALSIFY spec screenshot wrote `Task(...)` imported from `pipelines.tasks.Task`. The **correct verified path is lowercase** `cognee.modules.pipelines.tasks.task` (class name `Task`). Use that. `get_graph_engine` is `await`-ed; `get_vector_engine` is **not**.
---
## 1. Core Mechanism
### 1.1 Node & edge vocabulary (see §2 for full model)
- Nodes: `InvestigationQuestion` (root), `Hypothesis`, `Evidence`, `Conclusion`.
- Edges (relationship_name string on `add_edge`):
- `depends_on` : **Conclusion → Evidence** — *THE forward-propagation rail.* A Conclusion `depends_on` the Evidence it rests on. `edge_properties = {"critical": bool}`.
- `supports` : **Evidence → Hypothesis** — evidence corroborates a hypothesis. `edge_properties = {"weight": float}`.
- `refutes` : **Evidence → Hypothesis** — evidence contradicts a hypothesis. `edge_properties = {"weight": float}`.
- `supersedes` : **Evidence(new) → Evidence(old)** — written when the new fact back-dates / overrides an old evidence node. `edge_properties = {"confidence": float}`.
### 1.2 Truth-state lifecycle
`truth_alignment` is a list written to the node via `set_node_truth_state`. FALSIFY uses these canonical single-element states:
```
alive — default; node participates in recall context
refuted — Evidence directly contradicted by a newer fact (entry point of a cascade)
superseded — Evidence/Hypothesis replaced by a newer competing node (still exists, demoted)
invalidated — Conclusion whose supporting Evidence chain was refuted (forward-cascade victim)
forgotten — orphaned dead-end scheduled for surgical delete (transient; node then removed)
```
Transition rules:
```
alive --(new fact contradicts this Evidence, LLM-confirmed)--> refuted
alive --(a competing Evidence supersedes it)---------------> superseded
Conclusion.alive --(any critical depends_on Evidence is refuted)--> invalidated
{refuted|invalidated} node with NO surviving alive dependent/consumer --> forgotten --> delete_nodes + delete_data_points
Hypothesis.alive --(its only supporting Evidence became refuted AND a rival Hypothesis has surviving support)--> superseded
losing Hypothesis' rival --(gains the strongest surviving support)--> stays alive, feedback_weight promoted
```
Every state write also bumps `truth_epoch` to a monotonically increasing integer (epoch = the memify run counter). Recall filters on `truth_alignment` containing `"alive"`.
### 1.3 Contradiction detection (two-gate, deterministic-first)
The detector runs inside the `propagate_refutation` extraction/enrichment task. **Two gates** to kill nondeterminism:
1. **Vector prefilter (cheap, deterministic).** `hits = await ve.search("Evidence_claim", query_text=NEW_FACT, limit=5, include_payload=True)`. Candidates = hits with **cosine distance `score < 0.35`** (topically related — same subject). This narrows the LLM to plausibly-conflicting evidence only. Contradictory claims sometimes embed far apart, so claims MUST be normalized to `subject + predicate` phrasing at ingest, and `limit` kept ≥5.
2. **LLM adjudication (semantic).** For each candidate, call `LLMGateway.acreate_structured_output` with a **skeptical-analyst** system prompt and `response_model=ContradictionJudgement` (below). Only a verdict of `contradicts` (or `supersedes`) with `confidence ≥ 0.6` triggers refutation. This distinguishes *genuine contradiction* ("report was back-dated") from *topical overlap* ("also mentions the report").
```python
class ContradictionJudgement(BaseModel):
relation: Literal["contradicts", "supersedes", "supports", "unrelated"]
confidence: float # 0..1
rationale: str
```
**`--demo` / `DEMO_MODE` override:** when set, the seeded contradiction's target evidence id is pinned (`REFUTED_ID` env / seed constant), so the cascade runs on real graph APIs even if the LLM judge stalls or the key is flaky. This is the non-negotiable demo safety net.
### 1.4 Forward refutation propagation (exact algorithm)
Entry: an Evidence node `E` confirmed `refuted` (§1.3).
```
1. set_node_truth_state({E.id: {"truth_alignment": ["refuted"], "truth_epoch": epoch}})
set_node_feedback_weights({E.id: 0.0})
2. FORWARD CLOSURE — reverse-BFS along depends_on (Conclusion --depends_on--> Evidence).
Seeds = [E.id]. Traverse edges of type ["depends_on", "supersedes"] via:
nodes, edges = await ge.get_neighborhood([E.id], depth=4, edge_types=["depends_on", "supports"])
Collect every Conclusion C where a `depends_on` edge points from C into the refuted set
(transitively, up to depth 4). Because depends_on is Conclusion→Evidence, the "dependents"
are the SOURCES of those edges. A Conclusion is invalidated iff at least one of its
`critical: true` depends_on edges targets a refuted/invalidated node.
3. For each invalidated Conclusion C:
set_node_truth_state({C.id: {"truth_alignment": ["invalidated"], "truth_epoch": epoch}})
set_node_feedback_weights({C.id: 0.0})
4. PROMOTE competing hypothesis (promote_competing_hypothesis task):
- Find the Hypothesis H_dead whose only supporting Evidence is now refuted.
set_node_truth_state({H_dead.id: {"truth_alignment": ["superseded"], "truth_epoch": epoch}})
- Among rival Hypotheses still holding ≥1 alive `supports` Evidence, pick the one with the
highest summed support weight (read edge {"weight"}). Promote it:
set_node_feedback_weights({H_win.id: <boosted>}) # stays alive; becomes new frontier
5. RECORD new fact + supersedes edge (add_data_points dual-write):
- Materialize NEW_FACT as an Evidence DataPoint, add via add_data_points (writes graph + vector).
- add_edge(new_evidence.id, E.id, "supersedes", {"confidence": verdict.confidence})
```
**Direction summary:** refutation flows *from Evidence up to the Conclusions that depend on it* by walking `depends_on` edges backward (Conclusion is the edge source). Hypotheses are re-scored via their `supports`/`refutes` edges. No forward walk ever crosses a `supersedes` edge into an already-superseded node (prevents loops — see §4).
### 1.5 forget() orphan conditions (surgical delete)
`forget_orphan_deadends()` runs after propagation and grows a **death set** by traversal:
```
A node is FORGOTTEN (hard-deleted from graph + vector) iff ALL hold:
(a) its truth_alignment is refuted OR invalidated (never alive/superseded — superseded nodes
are kept as provenance), AND
(b) it has NO surviving consumer: no alive node reaches it via depends_on/supports
(checked with get_connections — _has_alive_alternative == False), AND
(c) it is not itself the target of a supersedes edge FROM an alive node (that node is the
provenance anchor of the new truth and must be retained).
Nodes that STILL feed a live node are retained even if refuted (partial refutation, §4).
```
Delete implementation (one shot per node batch):
```python
await ge.delete_nodes([str(id) for id in death_set])
await ve.delete_data_points("Evidence_claim", [evidence_ids_in_death_set])
await ve.delete_data_points("Conclusion_statement", [conclusion_ids_in_death_set])
```
> Provenance is kept: only *truly orphaned* dead-ends are hard-deleted. `refuted`/`superseded` nodes that still explain *why* the graph changed remain, carrying their state flag.
### 1.6 Cross-session persistence (the proof)
Truth-state is stored **on the node** in the graph DB (Ladybug/SQLite-backed), so it survives process restart. Session 2's `recall()` reads only `truth_alignment == alive` context. A parallel **plain-RAG baseline** (`recall(query_type=SearchType.RAG_COMPLETION)` or direct `ve.search`) does NOT read truth-state and re-cites the deleted/stale fact — this A/B is the scoreboard.
---
## 2. Data Model
### 2.1 Node DataPoint subclasses (`memory_core/models.py`)
| Class | Embeddable field | Other fields | Collection (auto) |
|---|---|---|---|
| `InvestigationQuestion` | `question: Annotated[str, Embeddable()]` | — | `InvestigationQuestion_question` |
| `Hypothesis` | `statement: Annotated[str, Embeddable(), Dedup()]` | `question_id: str`, `status: str="alive"`, `prior: float` | `Hypothesis_statement` |
| `Evidence` | `claim: Annotated[str, Embeddable(), Dedup()]` | `source_id: str`, `quote: str`, `stance: str` (`supports`/`refutes`), `asserted_at: str` (ISO date) | `Evidence_claim` (**the refutation entry point / vector prefilter target**) |
| `Conclusion` | `statement: Annotated[str, Embeddable()]` | `confidence: float`, `depends_on_ids: list[str]` | `Conclusion_statement` |
The collection name is `"{ClassName}_{embeddable_field}"` — auto-created on write. `Evidence_claim` is the collection the contradiction prefilter searches.
### 2.2 Edge types (already listed §1.1)
`depends_on` (Conclusion→Evidence, `{critical:bool}`), `supports`/`refutes` (Evidence→Hypothesis, `{weight:float}`), `supersedes` (Evidence→Evidence, `{confidence:float}`).
### 2.3 Node-level metadata / state fields (written via engine APIs, not model fields)
| Field | Written by | Meaning |
|---|---|---|
| `truth_alignment: list[str]` | `set_node_truth_state` | lifecycle state (§1.2). Recall filter key. |
| `truth_epoch: int` | `set_node_truth_state` | monotonically increasing revision epoch. |
| `feedback_weight: float` | `set_node_feedback_weights` | confidence/health (0.0 = dead, promoted hypotheses boosted). |
| `asserted_at` (model field on Evidence) | ingest | timestamp used for supersede tie-breaks (`newer = max(asserted_at)`). |
| `source_id` (model field) | ingest | provenance handle for the source document. |
### 2.4 Dedup across sessions
`Hypothesis` and `Evidence` set `identity_fields` in `metadata` (e.g. `["question_id","statement"]`). DataPoint generates a stable identity id from these (`DataPoint.py:76-81`), so re-adding the same belief in a later session updates the existing node instead of duplicating it — essential for cross-session refutation to land on the right node.
---
## 3. User / Demo Flow
### 3.1 End-to-end story (money shot)
> **Question:** *"Did Company X know about the defect before the recall?"*
>
> **Session 1** builds:
> - Hypothesis **A**: "X knew via QA report, Mar 2021" (supported by Evidence `E_qa`)
> - Hypothesis **B**: "X knew via supplier email, Jan 2021" (supported by Evidence `E_email`)
> - Hypothesis **C**: "X didn't know" (unsupported)
> - Conclusion **K**: "X knew by March 2021" — `depends_on(K → E_qa, critical=True)`
>
> Saved graph persists (survives restart).
>
> **Session 2** (reopened next day): analyst pastes ONE line — *"Forensic audit: the March QA report was back-dated."*
> 1. `Evidence_claim` vector prefilter finds `E_qa` (distance < 0.35).
> 2. LLM judge: `contradicts`, confidence 0.9.
> 3. `E_qa → refuted`; **forward BFS** over `depends_on` finds **K** (critical dep) → `K invalidated`.
> 4. **A** loses its only support → `A superseded`; **B** ignites as the new frontier (feedback_weight promoted) — in ~3s.
> 5. `forget_orphan_deadends`: K is orphaned (no alive consumer) → hard-deleted from graph + vector. `E_qa` kept as `refuted` provenance (target of new fact's `supersedes` edge).
> 6. **Scoreboard:** FALSIFY recall now answers via **B (Jan 2021)**; the **plain-RAG baseline still cites the back-dated March QA report.**
### 3.2 Screen beat (≤30s, one screen)
Force-graph shows A/B/C + K. Analyst pastes the fact → **E_qa flashes red → K crosses out and vanishes → A dims (superseded) → B glows green as new frontier.** A live scoreboard panel: `FALSIFY: "B — Jan 2021 (supplier email)"` vs `Plain RAG: "March 2021 QA report"` — captioned *"AI revised, not forgot."*
### 3.3 Cross-session proof
Restart the process (or run `main.py` a second time with `--session 2`). Because truth-state is on-node persisted, session 2's `recall()` never sees K or A-as-truth. Show the two recalls side by side.
---
## 4. Success Criteria
### 4.1 Unit-test assertions (pytest)
```
test_direct_refutation:
after propagate_refutation(new_fact contradicting E_qa):
assert get_node_truth_state([E_qa])[E_qa]["truth_alignment"] == ["refuted"]
test_forward_cascade_invalidates_conclusion:
assert get_node_truth_state([K])[K]["truth_alignment"] == ["invalidated"]
test_competing_hypothesis_promoted:
assert A.truth_alignment == ["superseded"]
assert "alive" in B.truth_alignment
assert feedback_weight(B) > feedback_weight(A)
test_orphan_forgotten_from_both_stores:
assert K.id NOT in (await ge has node) # gone from graph
assert K.id NOT in ve.search("Conclusion_statement", ...) # gone from vector
test_provenance_kept:
assert E_qa still exists with truth_alignment == ["refuted"] # NOT deleted (supersede anchor)
test_cross_session_recall_skips_dead:
recall("did X know?") context contains B, does NOT contain K or A-as-truth
test_baseline_still_stale:
RAG_COMPLETION / raw ve.search STILL returns the March QA claim # proves the differentiator
```
### 4.2 Expected `python main.py` output (judges, ~2 min)
```
[FALSIFY] Session 1: built belief graph for "Did Company X know...?"
Hypotheses: A(QA Mar'21) B(email Jan'21) C(didn't know)
Conclusion K depends_on E_qa
[FALSIFY] Session 2: new fact -> "March QA report was back-dated"
contradiction: E_qa (judge=contradicts conf=0.90)
forward cascade: K invalidated
A superseded -> B ignites (new frontier)
forgot 1 orphan (K) from graph + vector
[SCOREBOARD]
FALSIFY recall : X knew by Jan 2021 (supplier email) [revised]
Plain-RAG : X knew by Mar 2021 (QA report) [STALE]
```
Graceful exit if no API key: print a clear message + how to set `LLM_API_KEY`, exit code **0**.
### 4.3 Edge cases (must be handled)
| Case | Required behavior |
|---|---|
| **Cycle** (A depends_on B depends_on A) | BFS tracks a `visited` set; never revisit. `supersedes` edges never traversed into already-superseded nodes. Termination guaranteed. |
| **Multiple contradictions** (new fact refutes 2 Evidence nodes) | Each refuted independently; union of their dependent Conclusions invalidated; single forget pass over the merged death set. |
| **Partial refutation** (Conclusion depends_on E_qa AND E_alt, both critical) | Conclusion invalidated only if a *critical* dep is refuted AND no alive critical alternative remains (`_has_alive_alternative`). If E_alt still alive & critical satisfied, Conclusion stays alive; E_qa refuted but **retained** (still feeds a live node → not orphaned). |
| **Diamond dependency** (K depends_on E1,E2; E1,E2 both depends-chain to refuted E0) | Deduplicate via visited set so K is invalidated once, not twice; forget counts each node once. Explicit unit test required. |
| **Non-critical dep refuted** | Conclusion's `confidence` decays but stays `alive` (only `critical:true` deps invalidate). |
| **Contradiction judge false-positive risk** | Two-gate (vector `<0.35` + LLM `≥0.6`) + `--demo` pin. |
---
## 5. Tech Stack & Module Layout
### 5.1 Cognee APIs used (all §0-verified)
- Persistence: `cognee.remember(session_id=...)`, `cognee.recall(session_id=..., query_type=...)`.
- Enrichment pipeline: `cognee.memify(extraction_tasks=[Task(collect_belief_subgraph)], enrichment_tasks=[Task(propagate_refutation), Task(promote_competing_hypothesis), Task(forget_orphan_deadends), Task(add_data_points)], data=[{}], node_type=NodeSet, node_name=[question_id])`.
- Graph: `get_graph_engine()` → `get_neighborhood` (depth=4, edge_types=["depends_on","supports"]), `get_connections`, `set_node_truth_state`, `get_node_truth_state`, `set_node_feedback_weights`, `add_edge`, `delete_nodes`.
- Vector: `get_vector_engine()` → `search("Evidence_claim", ...)`, `delete_data_points`.
- LLM: `LLMGateway.acreate_structured_output(text_input, system_prompt, response_model=ContradictionJudgement)`.
- Storage: `add_data_points` task (`cognee/tasks/storage/add_data_points.py:31`) for dual graph+vector write of the new fact.
- Baseline: `SearchType.RAG_COMPLETION` (from `cognee/modules/search/types/SearchType.py`).
### 5.2 Module layout (write under `/workspaces/hackathon-app/`)
```
/workspaces/hackathon-app/
main.py # `python main.py` entry — runs seed + demo + scoreboard; graceful no-key exit 0
requirements.txt / pyproject # deps: cognee (editable ref) + minimal
.env.template # LLM_API_KEY, LLM_PROVIDER=openai, LLM_ENDPOINT (custom endpoint), LLM_MODEL, DEMO_MODE
memory_core/
__init__.py
models.py # DataPoint subclasses (§2.1) + ContradictionJudgement
edges.py # edge-name constants: DEPENDS_ON, SUPPORTS, REFUTES, SUPERSEDES
tasks.py # collect_belief_subgraph, propagate_refutation,
# promote_competing_hypothesis, forget_orphan_deadends
falsify.py # orchestration: build_graph(), revise(new_fact), scoreboard()
seed.py # demo corpus: Company-X recall investigation (A/B/C + K)
ui/
graph.html # react-force-graph (color by truth_state, red-flash-then-remove ripple ≤3s)
server.py # tiny static+JSON server feeding get_neighborhood snapshots
tests/
test_propagation.py # §4.1 assertions
test_edge_cases.py # cycle / diamond / partial / multi
```
### 5.3 Visualization approach
- **react-force-graph** (CDN, single `graph.html`) reads a JSON snapshot built from `ge.get_neighborhood([question_id], depth=4)`.
- Node color keyed on `truth_alignment`: alive=green, refuted=red, invalidated=grey-strikethrough, superseded=dim-amber. Forgotten nodes: **red-flash animation then removed** from the sim (≤3s ripple).
- Scoreboard panel overlays FALSIFY-vs-RAG answers. No external services — server is stdlib/`http.server` or FastAPI already in cognee.
### 5.4 Config / runtime constraints
- OpenAI-compatible: honor `LLM_PROVIDER` (`openai` or `custom`), `LLM_ENDPOINT`, `LLM_MODEL`, `LLM_API_KEY`. Defaults: LanceDB (vector) + Ladybug (graph) + SQLite (relational) — **zero external services**.
- `DEMO_MODE=1` (or `--demo`) pins `REFUTED_ID` so the cascade+forget run on real APIs regardless of LLM flakiness.
- Never modify `/workspaces/cognee`. All writes under `/workspaces/hackathon-app/` (absolute paths).
---
## 6. Non-negotiables (contract invariants)
1. Forward propagation uses `get_neighborhood` + `set_node_truth_state` (on-node, persistent). No in-memory-only state.
2. Surgical forget deletes from **both** graph (`delete_nodes`) and vector (`delete_data_points`); provenance (`refuted`/`superseded`) nodes are retained.
3. Cross-session persistence of disbelief is demonstrated (restart / session 2 recall skips dead branches).
4. The A/B scoreboard (FALSIFY revised vs plain-RAG stale) is shown every run.
5. Contradiction detection is two-gate (vector prefilter `<0.35` + LLM judge `≥0.6`) with a `--demo` deterministic override.
6. `Task` imported from `cognee.modules.pipelines.tasks.task` (lowercase); `get_graph_engine` awaited, `get_vector_engine` not.
|