Back to Knowledge Hub
Platform

Recover a Browser Session After a Network Interruption

Use bounded retries, checkpoints, and privacy-aware ownership to recover browser work after a network interruption without duplicating side effects.

BotBrowser Team

Documentation

Want the structured docs for Platform?

This article lives in the editorial library. For step-by-step setup, reference material, and ongoing updates, jump into the docs section.

A network interruption does not tell you whether a browser action completed. The connection may fail before a request, after the server accepted it, or while the worker was waiting for the response. Recover safely by preserving the session assignment, recording a checkpoint, and retrying only an operation whose completion can be confirmed.

TL;DR

  • Classify the failed channel: page request, websocket, proxy route, or control plane.
  • Keep profile, storage, locale, and route ownership stable during a recovery generation.
  • Retry reads and idempotent operations with a bounded budget; confirm writes before repeating them.
  • Treat browser-visible cleanup as different from server revocation and host-level deletion.
  • BotBrowser can repeat an authorized recovery checkpoint, but it cannot infer whether an external side effect was accepted.

Contents

A browser session moves from a checkpoint through a bounded network retry to a verified or quarantined result.

Freeze the session contract

Before opening a page, record a job reference, approved profile, storage owner, browser release, route policy, and lease generation. This record is an operational index; it must not contain passwords, page bodies, or unnecessary URLs. Keep one owner for the lease and reject messages from an older generation after a repair begins.

Separate durable intent from volatile state. Durable intent says what action is expected, the last accepted checkpoint, the retry budget, and the cleanup rule. Open pages, renderer memory, websocket connections, and in-flight requests are volatile. Do not pretend that a browser process restart reconstructed them.

Use checkpoints and a decision table

A checkpoint is an application-defined acceptance boundary: a read response was validated, a document was stored in an approved destination, or a business transaction returned a documented acknowledgement. A screenshot or successful navigation alone is not proof that a side effect completed.

Observed boundarySafe first actionRetry ruleExpected result
Read failed before responseReconnect and repeat the readBounded retry with jitterSame validated read or a classified failure
Write timed out with no acknowledgementQuery the application statusRepeat only with an idempotency keyOne recorded outcome, not two writes
Control channel lostPause admission and reconcile lease generationDo not replay page actions blindlyOne owner resumes or the job is quarantined
Route policy unavailableKeep the assignment pausedDo not switch to an unrelated routeA visible recoverable failure

Copyable recovery fixture:

const checkpoint = { version: 12, generation: 4, action: 'save-report' };
const outcome = await app.lookupByIdempotencyKey(checkpoint.action, jobKey);
if (outcome.confirmed) return outcome;
if (checkpoint.version !== (await lease.version())) throw new Error('reconcile ownership');
return app.saveReport({ jobKey, idempotencyKey: jobKey });

The fixture illustrates a rule, not a universal API: the application must provide a status lookup or idempotency contract. Without one, stop at the checkpoint and ask an owner to reconcile.

Recover with a bounded network budget

Classify whether the interruption affected a page request, a websocket, the proxy route, or the worker control channel. These failures have different recovery actions. Use a maximum attempt count and maximum age, with jittered delays. Rapid unlimited loops can duplicate writes and overload an unhealthy upstream service.

Keep retries inside the declared route and regional policy. A route change can change the session's network identity; close the old lease and create a new one if policy requires that change. Record restore generation, checkpoint version, attempt count, and result category without logging tokens or full private URLs.

When the budget is exhausted, mark the job recoverable-failure, preserve the last accepted checkpoint, drain the context, and quarantine uncertain storage. Releasing capacity before cleanup makes a later worker inherit ambiguous state.

The retry budget should be tied to the job's purpose, not chosen as a universal number. A one-time report download may allow two short reconnects, while a financial update may allow no automatic replay after an unknown acknowledgement. Set both an attempt limit and a wall-clock limit so a process cannot keep an old lease forever. Include the reason for each attempt in the recovery record, such as read-timeout, socket-closed, or status-reconciliation. This makes an operator's decision reproducible without retaining the entire browsing transcript.

Consider a form submission that times out immediately after the submit button is pressed. The page may show no change, the server may have accepted the form, or a validation response may have been lost. The safe sequence is to stop submitting, query the application's status using the operation identifier, and inspect the returned revision or receipt. If the service reports no matching operation and documents that lookup as authoritative, a single retry may be allowed. If lookup is unavailable or its result is stale, leave the operation in an unknown state. A new page load can provide evidence, but it does not make a second submission safe by itself.

For a download, preserve the destination as an untrusted temporary artifact until its length, digest, or application-level manifest is checked. A completed network stream can still produce a truncated file if the connection ended before the final bytes. Keep the temporary file separate from the path that downstream users consume, and publish it only after verification. If the browser resumes a partial transfer, record the byte range and server response that authorized the resume. Do not silently append bytes from a different URL, account, or route.

For a long-lived websocket, the browser may have received events that were not committed to local application state when the socket closed. On reconnect, ask the service for the last acknowledged revision and request a bounded delta or a fresh snapshot. Apply events in revision order and reject duplicates. If the service cannot provide an ordered view, discard the volatile in-memory queue and start from the last durable checkpoint. This preserves a clear boundary between what the browser displayed and what the service considers accepted.

Protect privacy and verify the result

Cookies, local storage, IndexedDB, cache entries, downloads, and server sessions have different owners and retention rules. A browser restart does not revoke a server session. MDN's Clear-Site-Data documentation describes origin-scoped cleanup, not deletion of remote records or unrelated host files. W3C privacy principles support purpose limitation and data minimization.

After recovery, compare structured evidence: assignment, route policy, checkpoint order, completion acknowledgement, and cleanup result. Use a fresh authorized context to verify that the expected application state is visible. Do not retain raw page content when a boolean assertion answers the question. If server revocation or a remote mutation is asynchronous, record it as pending rather than claiming completion.

Privacy review should cover the recovery record as well as the browser profile. A job identifier can become identifying when combined with a route, timestamp, and account reference, so keep those fields scoped to the smallest team that needs them. Prefer opaque identifiers over email addresses or account names. Store diagnostic screenshots only when a visual fact cannot be represented as structured data, and crop or redact before upload. A screenshot that proves a button was disabled may still reveal unrelated messages, names, or account balances.

Standards help define these boundaries, but they do not supply an application recovery policy. HTTP status codes describe the response that reached the client, not whether a business transaction was committed. The Fetch API can reject a request when a network failure occurs, yet that rejection does not prove that the server saw no bytes. Clear-Site-Data applies to an origin's browser-managed stores and is not a general remote deletion command. Use these standards as evidence about the browser and transport layer, then obtain separate evidence from the service that owns the record.

A useful acceptance record has one row per authority: browser, transport, application, and storage. The browser row can say whether the context closed and whether origin cleanup was requested. The transport row can identify the route class and whether a response was received. The application row can cite a receipt, revision, or status lookup. The storage row can identify whether a verified artifact was published or quarantined. Keeping these rows separate prevents a successful browser cleanup from being mistaken for account revocation or host-level deletion.

BotBrowser capability and limitation

BotBrowser can repeat an authorized browser-session checkpoint in a declared profile, release, route, and context, and compare visible outcomes after a controlled network interruption. It cannot determine whether an external server accepted an unacknowledged write, revoke server sessions, inspect every host artifact, or guarantee uninterrupted runtime. The application owner remains responsible for idempotency, authorization, retention, and final acceptance.

Use BotBrowser as the controlled browser participant in that contract. Supply a profile and route that the owner has approved, define the checkpoint before the run, and make the expected result observable without exposing secrets. A useful fixture might open a report, save the last known revision, interrupt the connection, reconnect within the same lease, and verify that the report is either available at that revision or explicitly marked pending. The fixture should also assert that no second report was created when the first save had an unknown outcome.

BotBrowser's visible result is strongest when paired with an application receipt. For example, it can confirm that a page displays revision 18, that a context was closed, or that a download was moved to quarantine. It cannot turn those observations into proof that a remote payment, deletion, message, or account change occurred. Keep the service's acknowledgement and BotBrowser's observation as two fields in the run record. This also makes failures easier to diagnose when the browser is healthy but the service is delayed, or when the service committed a change but the browser lost the response.

Do not use a recovery run to broaden authorization. A reconnect should retain the same purpose, profile scope, and data access as the interrupted run. If the user changes the requested action, account, destination, or retention rule, create a new authorized operation with a new identifier and checkpoint chain. BotBrowser can help execute that new operation, but it cannot approve the change or infer that the earlier lease should be reused.

Related reading: Long-running browser context resilience and Network layer browser consistency.

Final thoughts

Recover the intent, not the lost connection. Freeze ownership, check the last accepted boundary, reconcile uncertain writes, and use a bounded retry budget. The next safe action is to run one authorized fixture with an injected network interruption and verify both the application result and cleanup record.

For a production runbook, keep the recovery record small: a stable job identifier, checkpoint version, attempt number, route class, and a result such as confirmed, unknown, or quarantined. Keep secrets, full URLs, request bodies, downloads, and screenshots under separate retention rules. A compact record is easier to compare across retries and less likely to become a second copy of private browsing data.

Each checkpoint should name the authority that can answer the next question. That authority may be an idempotency-key lookup, a transaction-status endpoint, an application acknowledgement, or an operator decision. If no authority can distinguish accepted from unknown, preserve the uncertainty, stop the lease, and make reconciliation the next action instead of guessing.

Browser evidence and service evidence are different. Browser evidence can show that a request was sent, a response was parsed, a context was closed, or a cache was cleared. Service evidence can show that a record was committed, deduplicated, or revoked. A passing browser assertion must not be rewritten as proof of a remote commit.

For an interrupted download, compare the expected length or application digest when available, quarantine ambiguous files, and make cleanup observable. For a disconnected websocket, reconcile the document revision and replay only operations the server marks safe. A new context provides a clean observation surface; it cannot recreate renderer memory.

Route changes deserve their own checkpoint because they can change network identity, location, policy, or authorization assumptions. Close the old lease, record why the route changed, and re-check policy before continuing. Do not use a route change to hide an unresolved write.

Final acceptance can be a short table of owner, checkpoint, server acknowledgement, browser cleanup, and retention decision. Mark each row observed, pending, or not applicable. Collect only signals needed for the next safe action, expire them on a defined schedule, and avoid combining fingerprint, proxy, and session identifiers into a durable identity record.

The same distinction helps when a browser is used for a long-running report, import, or review. Divide the work into checkpoints that a person can understand: input accepted, validation complete, remote operation acknowledged, local artifact verified, and cleanup complete. Each checkpoint should be monotonic. A later observation may add information, but it should not silently rewrite an earlier unknown result as a success. If the application needs to correct a result, record a new correction event and keep the original evidence available for audit.

Use an operation identifier that is stable across a reconnect but different for each user-authorized action. A retry counter alone is not an idempotency key because two workers can start with the same counter. The identifier should be scoped to the application and protected from accidental reuse. When the server exposes a lookup, ask for the operation identifier before sending the action again. When it does not, return an explicit reconciliation state to the caller rather than making the browser guess.

A timeout is also a data-quality event. It can leave the browser with a response that the page has not rendered, a rendered page whose server commit is unknown, or a server commit whose acknowledgement was lost. Capture only the smallest observation needed to distinguish those cases. For example, a response status and application revision may be enough; a full HTML capture is often unnecessary and can expose account data. Redaction should happen before logs leave the worker, not after a shared log has already been copied.

Cleanup should follow ownership, not convenience. The browser owns some client stores, the application owns account sessions and records, and the host or storage service owns files and backups. A recovery routine can close a page or context and can request origin-scoped browser cleanup where the browser supports it. It cannot promise that a remote record, another device, a backup, or an unrelated host directory has disappeared. Keep those claims separate in both the user interface and the evidence record.

When a worker is replaced, transfer the lease deliberately. Publish the checkpoint version, owner generation, and expiry; reject late messages from the previous owner; and make the replacement read the last accepted state before it sends anything. If the replacement cannot establish continuity, quarantine the work. This is slower than an unconditional replay, but it prevents two workers from acting on the same uncertain state.

Network interruptions can expose privacy problems that were already present. A retry may send a token twice, a proxy may record a destination, and a diagnostic screenshot may preserve a private page. Limit retries, use transport protections required by the application, redact logs, and set retention periods before the incident occurs. Recovery should reduce uncertainty without creating a larger copy of the user's data.

For testing, inject failures at named boundaries rather than switching off the whole network. Test a request before its response, a response before its acknowledgement, a websocket after a revision, and a route change before a retry. The expected result for each fixture should say whether a retry is allowed, what evidence is required, and what cleanup must happen. A passing fixture demonstrates the application's contract; it does not establish behavior for every browser, device, proxy, or server.

Review the recovery design with the people who own the remote side effect and the data retention policy. Browser automation can make a repeatable observation, but it cannot decide whether a purchase, deletion, message, or account change is authorized. Put that authorization decision outside the retry loop, and make the user-visible state explain whether the operation completed, is pending, or needs attention.

FAQ

Should every timeout be retried?

No. Retry a read or an operation with an explicit idempotency contract. For an unknown write outcome, query status first; otherwise quarantine it for review.

Can a new browser context continue the old session?

Only when policy permits the same profile and storage and their integrity is known. A new context is not evidence that volatile page state survived.

Does clearing browser data revoke the server session?

No. Browser cleanup and server-side revocation are separate boundaries and need separate evidence.

What does BotBrowser verify?

It can replay the declared, authorized browser checkpoint and compare visible results. It does not certify remote side effects, host deletion, or identity.

Sources

#Browser Sessions#Network Interruption#Recovery#Data Consistency#Privacy#BotBrowser

Take BotBrowser from research to production

The guides cover the model first, then move into cross-platform validation, isolated contexts, and scale-ready browser deployment.