Found by driving the core against a real local server rather than an injected
delivery stub, which is exactly where it could hide: fetch only rejects on a
network-level failure, so a 500, a 503 or a 429 resolved normally and the batch
was counted as delivered and dropped. The unit tests could not catch it because
their stub throws, and real fetch does not.
That is the likelier outage than a refused connection, so the retry added in the
previous commit was covering the rarer half of the problem.
Any non-2xx now throws and takes the retry path. Matching the Python core, which
retries every HTTP error rather than classifying them: the backoff and the queue
bound contain a payload that will never be accepted, because the re-queued batch
sits at the front and is the first thing evicted.
Two tests against the real default delivery path, stubbing fetch rather than the
delivery hook, so a 503 keeps the batch and a 200 clears it.
End to end against a local server, real fetch and real retry timing: events
arrive and carry a uuid, a 503 keeps the batch, an immediate retry is suppressed
by the backoff, the retained event is delivered on recovery with its original
uuid and no duplicate, and a refused connection behaves the same way. Nine of
nine.
Identity checked on the same path: opencode's project_hash is salted, differs per
account, is stable within one, and no raw project id appears in the payload.
openclaw emits two distinct identities across a key change, which is the defect
this branch fixes, observed on the wire rather than through a mock.
agent-plugin-core/ts 32, openclaw 431, opencode 35, pi-agent 89, deepseek 47,
Python core 242 unaffected.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
The Python plugin telemetry was hardened across #7322 to #7326. The TypeScript
side has the same defect classes and was not touched, because the two share no
code: agent-plugin-core/python generates into six bundles, agent-plugin-core/
typescript is a separate core each plugin wraps. Fixing one surfaced nothing
about the other, which is how this survived.
Three fixes, all confirmed by running the code rather than reading it.
A failed delivery deleted the batch. The core detached the queue before the
await and swallowed the error, so one blip destroyed the events with nothing
recording that it happened. Probed: two events in, delivery throws, queue goes to
zero, no retry ever. The batch is now put back, bounded by maxQueueSize and
biased to the newest so a long outage costs the oldest events rather than
unbounded memory, and repeated failures back off to a ceiling instead of retrying
every flush against a host that is blocking us. Every event now carries a uuid
stamped at capture, which is what makes the retry safe: PostHog collapses
anything it already accepted.
Deliberately no disk spool, and that is written into the code so it reads as a
decision. Python spools because its hooks are per-tool-call processes that exit
immediately. These plugins live inside a host for a whole session, so
re-queueing covers the same transient failures without the claim and lease
machinery that took three review rounds to get right on the Python side. What it
leaves uncovered is narrow: a session that both starts and ends offline.
openclaw used a cached email forever. The refresh was guarded by !hasEmail, so
after an API key change every event kept reporting under the previous account.
The email is now bound to a fingerprint of the key it was resolved for and only
used while those agree; a mismatch forgets the account and re-resolves. The
resolution latch is per key rather than once per process, so a key changed
mid-session is actually looked up. A row with an email and no fingerprint, which
is what an upgrade from the current version looks like, is verified rather than
adopted, matching the decision reached on #7325.
opencode hashed the project id unsalted, which is reversible for anyone who can
enumerate project ids. Salted with the API key rather than a stored per-install
value: it is already in play, it is high entropy, and it needs no new file and so
no write race to get wrong. Per account rather than per machine, which also keeps
joins working across machines, and it resets on key rotation consistently with
distinctId, which already did.
Also wired opencode's tests into CI. They existed and nothing ran them, so the
regression test asked for on #7322 would not have gated anything.
Verified: agent-plugin-core/ts 30, openclaw 431, opencode 35, pi-agent 89,
deepseek 47. The new tests were each checked against the unfixed code first; the
core ones fail 3 of 3 and the openclaw one fails without the fingerprint gate.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb