Review found three ways the first cut produced worse data than no salt at all.
All three came from keeping the salt as a key in the identity dict and doing an
unlocked read-modify-write.
Hooks are short-lived separate processes firing on every tool call, and people
run more than one agent window, so several processes would read {}, each mint
its own uuid4, and each hash with it. One repository hashed several ways in the
window before a writer won.
resolve_distinct_id holds a copy of that same dict across a network call to
/v1/ping/ with a 5s timeout, so whichever write landed second erased the other's
key: losing the salt changes repo_hash mid-stream, losing the email fires a
second $identify and splits the person.
_write_identity swallows OSError, and nothing memoized, so on a read-only or
full data directory every single event got a brand-new random salt — unbounded
cardinality in PostHog, which is strictly worse than the unsalted value it
replaced.
The salt now lives in its own file claimed with O_CREAT|O_EXCL, so exactly one
process wins and the losers read the winner's value, and it is memoized per
process. When it cannot be persisted the fallback is derived from the data
directory path: stable for the machine rather than random per call.
Its own file also means record() no longer creates telemetry-identity.json as a
side effect. is_first_run keys off that file, so the first cut would have
silently suppressed the install event — a production metric change hidden in a
docs PR.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
Two defects, one cause: the spool protocol infers ownership instead of holding
it, and never records progress.
Duplicate delivery after a partial failure. flush() posts the claim in batches of
100 and returns on the first failure, keeping the whole file. The retry then
posts every batch again, including the ones that already arrived — 150 recorded
events were delivered 250 times. Progress is now written back to the claim after
each successful batch, so a retry resumes where the send stopped and a crash
repeats at most one batch.
Duplicate delivery when two senders overlap. spool.replace(claim) is os.rename,
which preserves mtime, so a claim created after a quiet minute inherited the
spool's last-write time and looked abandoned the instant it existed. A second
sender starting while the first was still posting took it over and sent it too —
most likely at session end, when the MCP server's exit sender and the SessionEnd
flush worker both drain. Claims are now touched at claim time, and the per-batch
rewrite doubles as a lease heartbeat. _post makes one attempt with SEND_TIMEOUT
and no retry, so a heartbeat lands well inside the 120s lease; a test asserts
that margin so adding a retry loop to _post cannot silently break it.
Parked batches starved. _claim_spool only looked at parked .sending files when
no spool existed, and because sessions keep recording there usually was one — so
a batch parked by a failed send waited until the 7-day expiry deleted it unsent,
despite its own presence being what starts the sender in the first place.
flush() now drains the live spool and then parked claims in the same run, oldest
first, bounded. Expiry applies only after a genuine retry has failed, with the
attempt count carried in the filename.
A sender that gives up releases its lease rather than heartbeating on the way
out, so the next run picks the batch up promptly instead of waiting a full stale
window for a batch nobody is working on. A failing send stops the run, so one
broken connection cannot burn every parked batch's attempt budget at once.
Two existing tests asserted the old lifecycle and are updated in place, each
with a comment saying what changed.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb