40287f6f04
Two review findings from @karthik-indla on this PR. _drain returned (0, True) on any read failure, so a claim nothing was posted from counted as fully delivered. flush() then carried on to the next claim as though this one had arrived, and the single signal that says the run went badly never fired. The two cases are now separated: undecodable content is still quarantined and reported delivered, because there is nothing left to send and the rest of the run should continue, while an OSError leaves the file exactly where it is and reports undelivered. Quarantining there would discard events over a transient filesystem error, and nothing ever re-globs .corrupt. Retries had no time backoff. _release_claim backdated straight to immediately-reclaimable, so two senders meeting one momentary failure could walk a batch from attempt 0 to the limit within seconds and discard it, when a retry a minute later would have delivered. Releases now carry a cooldown that grows with the attempts already spent, clamped so the mtime never lands in the future and reads as a live lease. Four tests: an unreadable batch is neither delivered nor quarantined, undecodable content still is quarantined so one torn file cannot block every later claim, and attempts cannot be burned without waiting. The expiry test now ages the file between flushes, which is the wall time a real retry waits. Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb