Long-Running Browser Context Resilience
Design browser contexts that recover from crashes, network interruptions, and maintenance while keeping storage ownership and validation evidence consistent.
Want the structured docs for Platform?
This article lives in the editorial library. For step-by-step setup, reference material, and ongoing updates, jump into the docs section.
Long-running browser work is different from a short script. A short script can start a context, complete a page action, and exit before an interrupted connection or a gradual resource problem becomes visible. A context that remains active for hours or days has to survive ordinary browser crashes, temporary network failures, expired leases, maintenance, and application restarts. It also has to preserve a clear relationship between its profile, storage, route policy, owner, and recorded results.
The goal is not to keep every context alive forever. The goal is to make continuation, recovery, and retirement predictable. A resilient design knows which state can be reconstructed, which state must be preserved, and when a new context is safer than a repair. That distinction protects privacy and makes a result reproducible without retaining more session data than the workload needs.
Treat the context as a stateful lease
A context should have one explicit owner and one lifecycle record. The record can contain a stable job reference, an approved profile reference, the storage policy, the route policy, the browser release, and timestamps for creation, checkpoint, activity, drain, and close. It should not contain page credentials, raw page content, or unrelated browsing data. The record is an operational index, not a copy of the session.
Model the lifecycle as states such as planned, starting, active, degraded, draining, closed, and needs-review. Keep the transitions idempotent. A supervisor may receive the same close notification twice, or a worker may reconnect after an acknowledgement was lost. Repeating a transition must not create a second context, release a lease twice, or attach storage to another owner.
The lease has a deadline and a renewal rule. Renewal should prove that the worker still owns the browser and can report health. It should not be extended solely because a page is busy. When renewal stops, the scheduler pauses new work for that context and begins reconciliation. This prevents a disconnected worker from silently retaining an assignment while another worker starts a conflicting repair.
Keep identity assignment stable for the entire active lifetime. Profile, locale, storage, and route choices belong to the lease created before the first page. A recovery attempt may recreate the same approved assignment, but it should not quietly substitute an unrelated profile or route because the original is inconvenient. If policy requires a new identity, close the old lease and start a new session with a new record.
Separate durable intent from volatile browser state
The browser process, renderer state, open pages, in-memory application state, and network connections are volatile. A service restart can lose them without losing the intent of the job. Durable intent should describe what the worker is expected to do, the last accepted checkpoint, the retry budget, and the cleanup policy. It should be small enough to inspect and safe to replay.
Use an explicit checkpoint boundary. A checkpoint might mean that navigation completed, a document was saved to an approved destination, or a business transaction received a confirmed response. Define the boundary in application terms. A screenshot in memory is not automatically proof that a downstream action completed, and a successful navigation is not proof that the page work was accepted.
Persist checkpoints atomically with the lease version. A worker reads version 12, performs the action, and writes version 13 only after the completion condition is met. If two workers attempt to write the same version, one update fails and the loser must reconcile instead of assuming it owns the continuation. This simple compare-and-set pattern avoids duplicate completion and stale recovery decisions.
Do not treat every page detail as recoverable state. Cookies, local storage, downloads, and application drafts may have different retention and privacy rules. Decide which items belong in persistent storage, which are intentionally ephemeral, and which must be recreated through an authorized application flow. A snapshot that captures more than the policy allows can create a larger privacy problem than the original interruption.
Make startup and restore deterministic
Startup should validate the assignment before opening the first page. Check that the browser release is approved for the profile, the storage location is owned by the lease, the route policy is available, and the worker has a compatible configuration. If a precondition fails, report an assignment error and keep the lease out of the active queue. Starting a partially configured context creates ambiguous evidence and makes a later retry harder to interpret.
Restore should use a fixed sequence. Reattach the approved profile and storage policy, create the context, apply the route and locale settings that belong to the assignment, and then run a lightweight health check. Only after that check passes should the worker resume application work. Record the restore attempt and its result, including whether the context was newly created or resumed. The log needs enough information to compare runs without exposing page data.
Use a restore generation to distinguish an old worker from a current one. When the supervisor decides that generation 4 is dead, it creates generation 5 and records the decision. Messages from generation 4 are rejected or routed to reconciliation. This is safer than trying to infer ownership from connection order, because a delayed message can arrive after a repair has already started.
Startup and restore should be bounded. A context that cannot become healthy within the measured startup window moves to needs-review or a controlled retry state. Do not keep extending the same attempt while it consumes a scheduler slot. Bounded waits make queue behavior visible and leave room for the supervisor to drain the affected browser group.
Handle browser crashes as a controlled transition
A crash is a lifecycle event, not a reason to replay every action blindly. First mark the lease as degraded and stop new work. Capture non-sensitive evidence such as the browser release, worker generation, last checkpoint version, context reference, and host health. Then decide whether the last checkpoint is sufficient to resume or whether the job must return to an application-defined recovery step.
If the browser group contains other authorized contexts, isolate the failed context in the scheduler record while the supervisor checks the group. A single context failure does not prove that every context is unsafe, but a process or host failure can affect all of them. The group owner should make that decision from health signals and lifecycle evidence, not from a caller's guess.
After a confirmed crash, create a new browser context only through the approved assignment path. Reuse the same profile and storage only when the policy says that the session may continue and the storage is known to be consistent. If storage integrity is uncertain, quarantine it for review and start with the permitted clean state. Do not mount one uncertain storage directory into two active contexts.
Recovery actions must be idempotent. A retry can encounter a page that already completed a side effect, a download that already exists, or a transaction that is awaiting confirmation. Use application acknowledgements, idempotency keys, or a read-after-write check where the application supports them. If no safe confirmation exists, stop at the checkpoint and require a review rather than guessing.
Design network retry as a budget
Network interruptions are common in long sessions. A lost connection can affect a page request, a websocket, a route endpoint, or the worker's own control channel. These are different failures and should not share one unlimited retry loop. Record the failure category, keep the original lease assignment, and use a bounded attempt budget with a delay that allows the network to recover.
Retry only operations that are safe to repeat. A read or a page reload may be retryable, while a payment, form submission, or file mutation may require an application acknowledgement before another attempt. The browser cannot determine that distinction from a timeout alone. The workload contract must specify the completion signal and the recovery action.
Keep retries inside the approved route and regional policy. A temporary route failure is not permission to switch to an unrelated route in order to keep a session moving. If a route change alters the intended session identity, end the current lease and create a new one through policy. This makes the resulting evidence honest: the record shows that the network assignment changed instead of presenting two environments as one continuous run.
Use jittered delays and a maximum retry age rather than a fixed rapid loop. Rapid retries can increase pressure while an upstream service or browser group is already unhealthy. When the budget is exhausted, mark the job as recoverable failure, preserve the checkpoint, and return capacity only after the context cleanup path completes. A final failure state is more useful than an apparently active job that cannot make progress.
Snapshot only what can be restored
A snapshot is useful when it represents a clear recovery boundary. It can include a checkpoint version, assignment metadata, a storage reference, a pending action type, and a timestamp. It should point to controlled storage rather than embedding credentials or page content in scheduler logs. Encrypt and retain any persistent data according to the application's policy.
Compare snapshot freshness with the lease deadline. A snapshot written after the worker lost ownership must not become the source for a new worker. Store the lease version and generation with the snapshot, then reject it when the ownership check fails. This prevents a delayed write from sending a repaired context back to an earlier point.
Snapshots should be forward compatible. Include a schema version and a migration rule, or explicitly reject snapshots from an unsupported browser or application release. A silent interpretation change can look like a successful restore while producing a different application state. When migration is not safe, complete the old lease through its supported close path and start a new session.
Keep a small retention window. Long-running work can generate many checkpoints, but keeping every intermediate page state increases storage and privacy exposure. Retain the latest usable checkpoint plus the evidence required to explain a recovery decision. Remove expired snapshots through the same owner-controlled cleanup process used for other session data.
Recycle contexts at a policy boundary
Recycling is part of resilience, not an admission that every context is faulty. A context can be healthy while its browser group needs maintenance, its application session reaches a natural boundary, or its storage policy requires closure. Define recycling triggers from observed health and workload rules, such as a completed job group, a release drain, an unrepairable state, or an ownership timeout.
Drain before recycling. Stop new work, let the current action reach its completion or cancellation boundary, write the final checkpoint, close pages, close the context, and release the lease only after cleanup is confirmed. If close does not complete, move the resource to reconciliation and keep it out of admission. Releasing a slot before cleanup is complete converts an uncertain resource into hidden capacity debt.
Do not recycle merely to rotate an identity. A continuous authorized session should retain its assignment until its stated work is complete. When policy requires a new assignment, make the boundary explicit in the run record and ensure old storage is not attached to the new context. This preserves consistency across the session and keeps privacy ownership clear.
Group recycling should be coordinated by one owner. The owner pauses admission, drains active contexts, confirms leases, replaces the browser group, and resumes from a lower admission level while health signals are checked. A browser binary should never be replaced underneath active contexts. Rollback should restore the complete approved pairing of browser release, profile support, and worker configuration.
Observe consistency, not just uptime
Uptime alone says little about whether a long-running context is trustworthy. Observe lifecycle consistency: one lease maps to one owner, one active context maps to one storage assignment, checkpoints advance in order, and close events eventually match every open event. These are relationships between records, not private page observations.
Track queue age, lease renewals, restore attempts, crash categories, network retry age, checkpoint lag, close duration, browser group replacement, and residual resource trends. Use stable references and redact account identifiers, credentials, profile contents, and page URLs when they are not required for operations. A dashboard should show whether recovery is progressing, not recreate the user's session.
Alert on sustained trends and mismatches. A single network retry may be normal. A growing count of contexts in degraded, close operations that take longer over time, or leases without matching workers deserves an admission pause and reconciliation. Pair each alert with a bounded action such as stop admission, drain a group, quarantine a storage reference, or restart from a known checkpoint.
Keep application failures separate from browser and capacity failures. A valid page response that rejects a business input is not the same as a browser crash. A worker that loses its control channel is not the same as an upstream timeout. Separate categories help operators choose a recovery step without changing identity or storage unnecessarily.
Build reproducible recovery verification
Verification should exercise the same lifecycle contract used in production. Start with an approved representative job and record its assignment, release, checkpoint sequence, and completion condition. Then introduce one controlled interruption at a time: stop the worker, interrupt the network path, close the browser group, or force a maintenance drain. The test should ask whether the service returns to a known state, not whether it can hide the interruption.
Compare the pre-interruption and post-recovery records. The profile assignment, locale policy, route policy, storage ownership, checkpoint ordering, and final application result should be explainable from the run record. A different result may be valid when the application changed, but the reason should be visible. Avoid comparing raw page captures when a small structured record answers the consistency question.
Repeat the verification on each supported browser release and deployment image that changes lifecycle behavior. Keep the workload, authorization, and completion criteria stable enough for comparison. Record the environment and test date, and separate a failed collection from a valid application result. A missing result is not evidence that the browser produced a different result.
Recovery tests should include cleanup. After a failed attempt, confirm that the old context is closed, the lease is no longer admitted, uncertain storage is isolated, and the replacement context has one owner. Then run a second approved job to prove that capacity returned without inheriting stale state. This catches leaks that a single successful recovery test can miss.
For a public view of supported browser consistency, visit the BotBrowser Proof Center. The browser context memory guide discusses resource planning, while scaling browser contexts covers admission limits and lifecycle ownership. These references complement a service's own authorized recovery records; they do not replace application acceptance criteria.
A practical operating checklist
Before enabling a long-running context, confirm:
- The lease has one owner, an expiry rule, and an idempotent state machine.
- Profile, storage, route, locale, and browser release are assigned before page work.
- Durable intent and checkpoints exclude secrets and unnecessary page data.
- Restore generations reject delayed messages from an old worker.
- Crash and network recovery have bounded budgets and application completion checks.
- Recycling drains work and confirms cleanup before releasing capacity.
- Snapshot schema, freshness, retention, and migration rules are explicit.
- Dashboards expose lifecycle relationships without retaining private session contents.
- Controlled interruption tests compare records and verify cleanup on every supported release.
During operation, pause admission when ownership, checkpoint, or cleanup relationships become unclear. Reconcile first, then decide whether to resume the existing assignment or start a new one. A short, explicit interruption is easier to explain than a context that appears healthy while its state is no longer trustworthy.
Long-running resilience is therefore a consistency discipline. The browser may restart, the route may recover, and the worker may be replaced, but the service should still be able to say which assignment was active, what work was confirmed, what state may be restored, and why a new context was created. That is the basis for privacy-aware recovery and reproducible browser automation.
Related Articles
Take BotBrowser from research to production
The guides cover the model first, then move into cross-platform validation, isolated contexts, and scale-ready browser deployment.