Commit Graph

3 Commits

Author SHA1 Message Date
Saket Aryan 44c6a07ddf fix(plugins): flush on exit regardless of backoff, and prefer the backlog over new events
Findings from an independent review of this branch.

The exit-time flush was gated by its own cooldown. beforeExit called the same
flush() that opens with a retryNotBefore check, so after any failed delivery a
process exiting inside the 2s to 60s window sent nothing and the queue died with
it. That is precisely the loss this branch exists to stop, and timer.unref makes
beforeExit often the only remaining chance. flush(force) now skips the cooldown
and beforeExit passes it.

Overflow kept the newest and evicted the batch being retried, which threw away
exactly the events the retry exists to save. Both the failure path and capture()
now keep the backlog and drop the new event instead, matching Python's record(),
which refuses new events once the spool is full.

That change needs a bound, so delivery now gives up after five attempts, as the
Python core does. Without one a payload the server will never accept would be
retried for the whole session and, with the backlog now preferred, would hold the
queue against everything behind it.

Events carry a capture-time timestamp. They now sit through backoff and across
an entire outage, so without one PostHog records them at whatever moment delivery
happened to succeed. It also matters for the uuid dedupe, whose key includes the
event date.

openclaw wiped a resolved account on any capture without an apiKey: the
fingerprint comparison was `undefined === ""`, so a keyless call looked like a key
change. Guarded on a real key being present. Masked today because every call site
supplies one, which is why only a test found it.

Two smaller ones. opencode's PostHog source was "plugin", which named no
particular plugin and matched no vocabulary; it is now OPENCODE_PLUGIN like every
other surface, and a saved insight filtering source = "plugin" needs repointing.
And an em dash in opencode's published description had been rewritten to a —
escape by my own json.dumps when adding the test script; restored.

One openclaw test asserted only not.toThrow() under a name claiming it checked
the identity, and under the fingerprint gate the path it exercised no longer uses
the email at all. It now pins the real condition.

core 35, openclaw 432, opencode 35, pi-agent 89, deepseek 47. The wire e2e still
passes 9 of 9 against a real local server, and the identity e2e 7 of 7.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-17 19:30:59 +05:30
Saket Aryan c05556f3cf fix(plugins): stop the TypeScript telemetry losing events and misattributing accounts
The Python plugin telemetry was hardened across #7322 to #7326. The TypeScript
side has the same defect classes and was not touched, because the two share no
code: agent-plugin-core/python generates into six bundles, agent-plugin-core/
typescript is a separate core each plugin wraps. Fixing one surfaced nothing
about the other, which is how this survived.

Three fixes, all confirmed by running the code rather than reading it.

A failed delivery deleted the batch. The core detached the queue before the
await and swallowed the error, so one blip destroyed the events with nothing
recording that it happened. Probed: two events in, delivery throws, queue goes to
zero, no retry ever. The batch is now put back, bounded by maxQueueSize and
biased to the newest so a long outage costs the oldest events rather than
unbounded memory, and repeated failures back off to a ceiling instead of retrying
every flush against a host that is blocking us. Every event now carries a uuid
stamped at capture, which is what makes the retry safe: PostHog collapses
anything it already accepted.

Deliberately no disk spool, and that is written into the code so it reads as a
decision. Python spools because its hooks are per-tool-call processes that exit
immediately. These plugins live inside a host for a whole session, so
re-queueing covers the same transient failures without the claim and lease
machinery that took three review rounds to get right on the Python side. What it
leaves uncovered is narrow: a session that both starts and ends offline.

openclaw used a cached email forever. The refresh was guarded by !hasEmail, so
after an API key change every event kept reporting under the previous account.
The email is now bound to a fingerprint of the key it was resolved for and only
used while those agree; a mismatch forgets the account and re-resolves. The
resolution latch is per key rather than once per process, so a key changed
mid-session is actually looked up. A row with an email and no fingerprint, which
is what an upgrade from the current version looks like, is verified rather than
adopted, matching the decision reached on #7325.

opencode hashed the project id unsalted, which is reversible for anyone who can
enumerate project ids. Salted with the API key rather than a stored per-install
value: it is already in play, it is high entropy, and it needs no new file and so
no write race to get wrong. Per account rather than per machine, which also keeps
joins working across machines, and it resets on key rotation consistently with
distinctId, which already did.

Also wired opencode's tests into CI. They existed and nothing ran them, so the
regression test asked for on #7322 would not have gated anything.

Verified: agent-plugin-core/ts 30, openclaw 431, opencode 35, pi-agent 89,
deepseek 47. The new tests were each checked against the unfixed code first; the
core ones fail 3 of 3 and the openclaw one fails without the fingerprint gate.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-17 18:43:27 +05:30
Kartik 73e7b8763a refactor(integrations): shared agent plugin runtimes and native adapters (#7203) 2026-09-08 23:32:25 +05:30