File size: 4,476 Bytes
d9ffd67
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
---
title: Seeder auxiliary Redis writes crash on a single Upstash timeout
date: 2026-07-17
category: database-issues
module: scripts/_seed-utils.mjs
problem_type: database_issue
component: background_job
symptoms:
  - "`FATAL: The operation was aborted due to timeout` after a successful fetch in seed-gdelt-intel"
  - "Railway badge flips red with PUBLISH_TIMEOUT class; 3 crashes in 25-run window, 0 successes"
  - "Auxiliary `writeExtraKey` SETs or `extendExistingTtl` EXPIRE pipelines time out, taking the whole run down"
root_cause: missing_tooling
resolution_type: code_fix
severity: medium
tags:
  - redis
  - upstash
  - seeder
  - retry
  - seed-utils
  - afterpublish
  - timeout
---

# Seeder auxiliary Redis writes crash on a single Upstash timeout

## Problem

`seed-gdelt-intel` was crashing with `FATAL: The operation was aborted due to timeout` during the post-fetch Redis write phase. The upstream GDELT fetch had already succeeded and the canonical key publish (`atomicPublish`) already retried transient failures, but the auxiliary timeline-key writes and TTL extensions in `afterPublish` used one-shot `fetch()` calls. A single Upstash latency spike turned a transient blip into a full seeder crash and a Railway "Deploy Crashed!" email.

## Symptoms

- `FATAL: The operation was aborted due to timeout` appears after `Extended TTL on N key(s)` / `WARNING: N key(s) were expired/missing` logs.
- The seeder diagnostic classifies the service as `PUBLISH_TIMEOUT` with a warning severity.
- The crash recurs (3 in the inspected window) because every run re-rolls the same dice against Upstash tail latency.

## What Didn't Work

- **Retrying only the canonical publish.** `atomicPublish` already wrapped its staging/canonical SET/DEL in `withRetry`, but that only protects the canonical key. The `afterPublish` auxiliary writes (`writeExtraKey`, `extendExistingTtl`) were left single-shot.
- **Catching the timeout inside `extendExistingTtl`.** That helper already caught errors and returned `false`, but `writeExtraKey` threw on any non-ok response or abort, and neither helper retried — so a transient timeout still failed the run.

## Solution

Wrap both auxiliary Redis helpers in the same retry contract already used by `redisCommand` and `atomicPublish`:

- `writeExtraKey` (`scripts/_seed-utils.mjs:668`) now wraps its SET call in `withRetry` with 2 retries and a 1s base delay.
- `extendExistingTtl` (`scripts/_seed-utils.mjs:722`) now wraps its `/pipeline` call in `withRetry` with the same budget.
- Permanent 4xx errors are tagged `nonRetryable` so they fail fast.
- HTTP 429 errors honor the upstream `Retry-After` header.
- 5xx, timeouts, and network tears retry with exponential backoff.

The boolean contract of `extendExistingTtl` is preserved: it still returns `true` only when every `EXPIRE` returns `1`. A successful response with some `EXPIRE` no-ops (missing/expired keys) is a real data condition, not a transient error, so it returns `false` without burning retries.

Fixed in PR [#5364](https://github.com/koala73/worldmonitor/pull/5364).

## Why This Works

The root cause was not a bad source or bad data — it was a missing resilience layer on the auxiliary write path. Upstash REST is served over the public internet; a single stalled request or brief 503 is expected at scale. The canonical publish path already treated these as retryable; the auxiliary path did not. Adding retry makes the failure mode symmetric across all Redis writes in a seeder run.

## Prevention

- When adding a new Redis helper in `scripts/_seed-utils.mjs`, decide its retry contract up front. Helpers that write seeded data should default to `withRetry` unless the caller explicitly needs fail-fast semantics.
- Keep error tagging consistent with `redisCommand`:
  - `PERMANENT_4XX_STATUSES``err.nonRetryable = true`
  - `429` → parse `Retry-After` into `err.retryAfterMs`
  - everything else (5xx, timeout, network tear) → let `withRetry` back off
- Add a regression test that fails the first call and succeeds on retry for any new Redis write helper. The existing tests for `writeExtraKey` and `extendExistingTtl` now cover timeout, 503, 429, and permanent 401 paths.

## Related Issues

- `diagnose-railway-seeders` skill class `PUBLISH_TIMEOUT` — post-fetch Redis publish timed out.
- Memory: [[feedback_never_memorize_a_workaround_for_a_tool_bug_fix_the_tool]] — fix the shared helper rather than working around it in one seeder.