Review finding. The comments described MAX_DELIVERY_ATTEMPTS as a per-batch
budget mirroring the Python core's. It is not: Python puts the attempt count in
the claim filename so it follows one batch, while this counter lives in the
closure and counts consecutive failed flushes, so events captured during an
outage join the same queue and are dropped with it.
Per-batch accounting would mean an attempt count on every event. The queue is
already bounded, so the simpler rule stands; the comments now describe it rather
than the Python one.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
Second review round. The first two are regressions from the first round.
The forced exit flush added fifteen seconds to host shutdown. Node re-emits
beforeExit whenever the handler schedules async work, so an unconditional
flush(true) looped until the five-attempt budget was spent, and against the real
3s delivery timeout that is 15s added to the shutdown of whatever editor or CLI
is hosting this. The backoff used to end that loop after one attempt; removing it
for the forced path removed the only thing bounding it. Measured at 15008ms, now
6001ms with a one-shot latch, and 12ms when delivery is healthy, which is the
only case most people ever see.
openclaw deleted a legacy account before it had anything to replace it with. An
install predating keyFingerprint has an email and no fingerprint, so the
comparison failed and clearResolvedAccount() ran immediately; if the re-resolve
then failed because the user was offline the email was gone from disk for good,
and the per-key latch was already set so nothing retried. Now cleared only when a
real fingerprint disagrees, which is the same trade the Python core makes and
documents: verify, and keep what you have until the verification succeeds. The
latch is released on a failed lookup so the next capture tries again.
The overflow test did not exercise the path it is named for: with only two
events captured the re-queue had an empty queue to merge into, so the slice on
the failure path never ran, and that is the half deciding which end gets dropped.
It now fills past the cap before flushing and asserts the exact surviving order.
core 36, openclaw 432, opencode 35, deepseek 47. Wire e2e 9 of 9 and identity
e2e 7 of 7 still pass against a real local server.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
Findings from an independent review of this branch.
The exit-time flush was gated by its own cooldown. beforeExit called the same
flush() that opens with a retryNotBefore check, so after any failed delivery a
process exiting inside the 2s to 60s window sent nothing and the queue died with
it. That is precisely the loss this branch exists to stop, and timer.unref makes
beforeExit often the only remaining chance. flush(force) now skips the cooldown
and beforeExit passes it.
Overflow kept the newest and evicted the batch being retried, which threw away
exactly the events the retry exists to save. Both the failure path and capture()
now keep the backlog and drop the new event instead, matching Python's record(),
which refuses new events once the spool is full.
That change needs a bound, so delivery now gives up after five attempts, as the
Python core does. Without one a payload the server will never accept would be
retried for the whole session and, with the backlog now preferred, would hold the
queue against everything behind it.
Events carry a capture-time timestamp. They now sit through backoff and across
an entire outage, so without one PostHog records them at whatever moment delivery
happened to succeed. It also matters for the uuid dedupe, whose key includes the
event date.
openclaw wiped a resolved account on any capture without an apiKey: the
fingerprint comparison was `undefined === ""`, so a keyless call looked like a key
change. Guarded on a real key being present. Masked today because every call site
supplies one, which is why only a test found it.
Two smaller ones. opencode's PostHog source was "plugin", which named no
particular plugin and matched no vocabulary; it is now OPENCODE_PLUGIN like every
other surface, and a saved insight filtering source = "plugin" needs repointing.
And an em dash in opencode's published description had been rewritten to a —
escape by my own json.dumps when adding the test script; restored.
One openclaw test asserted only not.toThrow() under a name claiming it checked
the identity, and under the fingerprint gate the path it exercised no longer uses
the email at all. It now pins the real condition.
core 35, openclaw 432, opencode 35, pi-agent 89, deepseek 47. The wire e2e still
passes 9 of 9 against a real local server, and the identity e2e 7 of 7.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
Found by driving the core against a real local server rather than an injected
delivery stub, which is exactly where it could hide: fetch only rejects on a
network-level failure, so a 500, a 503 or a 429 resolved normally and the batch
was counted as delivered and dropped. The unit tests could not catch it because
their stub throws, and real fetch does not.
That is the likelier outage than a refused connection, so the retry added in the
previous commit was covering the rarer half of the problem.
Any non-2xx now throws and takes the retry path. Matching the Python core, which
retries every HTTP error rather than classifying them: the backoff and the queue
bound contain a payload that will never be accepted, because the re-queued batch
sits at the front and is the first thing evicted.
Two tests against the real default delivery path, stubbing fetch rather than the
delivery hook, so a 503 keeps the batch and a 200 clears it.
End to end against a local server, real fetch and real retry timing: events
arrive and carry a uuid, a 503 keeps the batch, an immediate retry is suppressed
by the backoff, the retained event is delivered on recovery with its original
uuid and no duplicate, and a refused connection behaves the same way. Nine of
nine.
Identity checked on the same path: opencode's project_hash is salted, differs per
account, is stable within one, and no raw project id appears in the payload.
openclaw emits two distinct identities across a key change, which is the defect
this branch fixes, observed on the wire rather than through a mock.
agent-plugin-core/ts 32, openclaw 431, opencode 35, pi-agent 89, deepseek 47,
Python core 242 unaffected.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
The Python plugin telemetry was hardened across #7322 to #7326. The TypeScript
side has the same defect classes and was not touched, because the two share no
code: agent-plugin-core/python generates into six bundles, agent-plugin-core/
typescript is a separate core each plugin wraps. Fixing one surfaced nothing
about the other, which is how this survived.
Three fixes, all confirmed by running the code rather than reading it.
A failed delivery deleted the batch. The core detached the queue before the
await and swallowed the error, so one blip destroyed the events with nothing
recording that it happened. Probed: two events in, delivery throws, queue goes to
zero, no retry ever. The batch is now put back, bounded by maxQueueSize and
biased to the newest so a long outage costs the oldest events rather than
unbounded memory, and repeated failures back off to a ceiling instead of retrying
every flush against a host that is blocking us. Every event now carries a uuid
stamped at capture, which is what makes the retry safe: PostHog collapses
anything it already accepted.
Deliberately no disk spool, and that is written into the code so it reads as a
decision. Python spools because its hooks are per-tool-call processes that exit
immediately. These plugins live inside a host for a whole session, so
re-queueing covers the same transient failures without the claim and lease
machinery that took three review rounds to get right on the Python side. What it
leaves uncovered is narrow: a session that both starts and ends offline.
openclaw used a cached email forever. The refresh was guarded by !hasEmail, so
after an API key change every event kept reporting under the previous account.
The email is now bound to a fingerprint of the key it was resolved for and only
used while those agree; a mismatch forgets the account and re-resolves. The
resolution latch is per key rather than once per process, so a key changed
mid-session is actually looked up. A row with an email and no fingerprint, which
is what an upgrade from the current version looks like, is verified rather than
adopted, matching the decision reached on #7325.
opencode hashed the project id unsalted, which is reversible for anyone who can
enumerate project ids. Salted with the API key rather than a stored per-install
value: it is already in play, it is high entropy, and it needs no new file and so
no write race to get wrong. Per account rather than per machine, which also keeps
joins working across machines, and it resets on key rotation consistently with
distinctId, which already did.
Also wired opencode's tests into CI. They existed and nothing ran them, so the
regression test asked for on #7322 would not have gated anything.
Verified: agent-plugin-core/ts 30, openclaw 431, opencode 35, pi-agent 89,
deepseek 47. The new tests were each checked against the unfixed code first; the
core ones fail 3 of 3 and the openclaw one fails without the fingerprint gate.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb