Retries that invent latency cliffs

A span that opens only after a silent retry can look like mass abandonment.

Code on a laptop screen beside a coffee cup

On mobile and flaky edges, a confirmation step that sits behind an automatic retry can silently fail for half the cohort. The chart shows a cliff; the application never asked those users to leave.

Before rewriting the flow, walk the journey on the networks your users actually use and watch the live stream. Note client-only fires versus server confirmation. Check identity continuity across hops.

A Latency Root-Cause Clinic stays narrow on purpose: one journey, deep QA, a checklist you can rerun after every release so the cliff does not return unnoticed.

All field notes Talk through a similar issue