Back to Knowledge Hub
Platform

Recovering from WebGPU Device Loss Without Losing the User

Design a bounded WebGPU recovery flow that preserves page state and prevents stale asynchronous work from replacing a newer GPU device.

BotBrowser Team

Documentation

Want the structured docs for Platform?

This article lives in the editorial library. For step-by-step setup, reference material, and ongoing updates, jump into the docs section.

When a WebGPU device is lost, the graphics path may stop working even though the rest of the page is healthy. The W3C specification exposes this lifecycle event through GPUDevice.lost; MDN describes its promise as pending for the device lifetime and resolving with loss information when the device is lost. Keep the user's task in ordinary application state, switch promptly to a usable non-GPU view, and treat another initialization as a new generation. A generation token is the key to ignoring late work from the device that has already been replaced.

A WebGPU device moves from ready to lost, the application keeps its regular view available, and a deliberate retry starts a new generation

Separate application state from GPU resources

The graphics device is a resource used to present or compute a result. It should not be the only place where the application keeps a document, selected item, form values, or other work the user expects to retain. Keep those values in normal page or application state. Then loss can remove the rendering resource without erasing the task itself.

GPUDevice.lost is a promise associated with one device. It settles when that device is lost; it is not a promise that a replacement will become available. The WebGPU specification also defines device loss as a change in device usability, not as a diagnosis of a physical component. Your interface should report the affected feature accurately, preserve available work, and avoid making claims about why a particular machine lost a device.

This is distinct from the privacy question in the WebGPU fingerprinting overview, which addresses how exposed signals can contribute to tracking. It is also different from browser release validation, which helps operators assess a browser/profile release pair and rollback plan. This guide answers a runtime application question: what should happen to an active feature after its current device stops being usable?

A generation-guarded recovery state machine

Use a monotonically increasing generation number whenever initialization is superseded or stopped. Every asynchronous continuation captures the generation that started it. Before it changes visible state, it must still match the current generation. The same check belongs in a device's lost handler: a late event from an old device must not replace the status of a newer one.

The following copyable controller demonstrates that boundary. The page-specific showFallback, showGraphics, announce, stopRenderLoop, detachGraphics, and disposeDeviceResources functions belong to the application. detachGraphics removes the previous surface from the active view; disposeDeviceResources releases its application-owned buffers and textures before device destruction. Application data remains outside this controller. initialize is called only when the user opens the feature or deliberately retries it. This sample does not retry automatically or enforce a retry counter; the feature owner applies any product retry budget.

let generation = 0;
let activeDevice = null;

function isCurrent(token) {
  return token === generation;
}

function stopGraphics() {
  generation += 1;
  const oldDevice = activeDevice;
  activeDevice = null;
  try {
    retireDevice(oldDevice);
  } finally {
    showFallback('Graphics stopped. Your page state is still available.');
  }
}

function retireDevice(device) {
  if (!device) return;
  try {
    stopRenderLoop(device);
  } finally {
    try {
      detachGraphics(device);
    } finally {
      try {
        disposeDeviceResources(device);
      } finally {
        device.destroy();
      }
    }
  }
}

async function initialize(gpu = navigator.gpu) {
  const token = ++generation;
  const oldDevice = activeDevice;
  activeDevice = null;
  showFallback('Starting the graphics feature…');

  try {
    retireDevice(oldDevice);
    if (!gpu) throw new Error('WebGPU is unavailable in this context.');
    const adapter = await gpu.requestAdapter();
    if (!isCurrent(token)) return;
    if (!adapter) throw new Error('No adapter was returned.');

    const device = await adapter.requestDevice();
    if (!isCurrent(token)) {
      retireDevice(device);
      return;
    }

    activeDevice = device;
    showGraphics(device);
    announce('Graphics are ready.');

    void device.lost.then(info => {
      if (!isCurrent(token)) return;
      generation += 1;
      activeDevice = null;
      try {
        retireDevice(device);
      } finally {
        showFallback(`Graphics stopped (${info.reason || 'device lost'}). Your page state is still available.`);
      }
    });
  } catch {
    if (isCurrent(token)) {
      const failedDevice = activeDevice;
      activeDevice = null;
      try {
        retireDevice(failedDevice);
      } finally {
        showFallback('Graphics could not start. Your page state is still available.');
      }
    }
  }
}

In production, handle the actual loss information according to the API contract and your support needs; do not turn it into a hardware classification. When an old requestDevice() resolves after stopGraphics() or a later initialize(), the token check destroys that obsolete device instead of showing it. A late rejection is also ignored. If the old lost promise settles after a new generation begins, its handler cannot overwrite the new view.

Choose a user-facing result

The transition should reflect what remains possible, not speculate about the cause of loss.

WebGPU is absent, the adapter is null, or initialization rejects. Keep the regular HTML view visible and say enhanced graphics could not start. Do not retry automatically; keep the feature optional.

A ready device's lost promise resolves. Mark only that generation as lost while keeping page data and controls available. Offer a retry control only when the user can reasonably choose to try again.

An active device is replaced and the new initialization fails or succeeds. Before requesting its replacement, stop and detach the old renderer, dispose its owned resources, and destroy its device. If the request fails, keep the fallback and do not restore the old device; if it succeeds, attach only the new device.

A newer generation starts while an older request is pending. Treat the old completion as stale and dispose any device it created. It must not change the new generation's status or view.

A deliberate retry succeeds. Attach the new device and render from current application state. Observe only the new device's loss promise.

A deliberate retry fails or the application retry budget is exhausted. Keep the fallback usable and explain that enhanced graphics remain unavailable. Stop retrying until a new user action or an application-defined reset; the budget is enforced by the feature owner, not by this sample.

If your product permits retry, make it a button or other explicit action with a clear disabled/loading state. A retry that runs in a tight loop can repeatedly allocate resources without improving the experience. Set a small application-owned budget or require a fresh user action after failure. Neither this sample nor GPUDevice.lost promises recovery; the feature owner must enforce any retry limit outside this controller.

Design the transition around ownership. The component that owns the task should decide whether the graphics surface is optional, what data must survive, and which view can continue without a device. A renderer can report that its device is unavailable, but it should not silently discard the document or independently launch a retry. Keeping this decision above the renderer gives the page one place to coordinate status, controls, and user intent.

Treat each initialization as a short transaction. First mark the previous generation obsolete, detach the old rendering path, and retain application data. Then request the resources for the new generation. Publish the new device only after the final generation check succeeds. If any asynchronous step fails, update the interface only when its token is still current. This ordering prevents a slow request from undoing a more recent user choice, such as closing the enhanced view or selecting a static representation.

The device-lost event and an application stop can race with cleanup. Invalidate the token before destroying or detaching the active device, so promise callbacks queued by cleanup cannot publish into the old generation. Make cleanup idempotent: repeated stop calls should leave the same ordinary view visible and should not destroy a replacement device accidentally. A renderer can keep its own cleanup routine, but the page-level generation remains the authority for which renderer is current.

Do not treat every loss reason as a recovery instruction. The loss information can help distinguish broad API-defined categories for a status message or a support case, but it does not identify a particular GPU or establish why the operating environment changed. The experience should remain safe when the reason is absent, unfamiliar, or not useful to the person. A concise message about the graphics feature is preferable to an unsupported hardware diagnosis.

Consider whether a feature can be reconstructed from ordinary state. If a canvas shows a chart derived from stored rows and filters, recreate the chart from those inputs after a new device is ready. If the GPU performed work whose intermediate values existed only in device memory, that work may not be recoverable; save checkpoints in application-owned storage when the user-visible task requires them. State explicitly which work is preserved and which operation must be repeated. Do not imply that destroy() saves or transfers device resources.

The page should also tolerate a device loss while a user action is in progress. Disable only controls that depend on the lost graphics path; keep navigation, cancel, export, and accessible alternatives available where they still make sense. If a pending action has not been committed to application state, either finish it using the non-GPU path or explain that it must be repeated. A status region announced to assistive technology can describe the transition without moving focus away from the user's current control.

An intentional call to GPUDevice.destroy() also ends that device's useful lifetime, so invalidate its generation before calling cleanup. Otherwise the resulting lost completion can look like an unexpected runtime failure and race with the page's stop transition. The token check makes both cases safe: a current device loss selects the fallback, while a loss caused by stopping an obsolete generation has no authority to change the current interface.

Avoid coupling recovery to page reload. Reloading may clear an in-progress form, reset navigation, repeat network requests, or make the person reconstruct their context. A device loss is a resource lifecycle event, not automatically a reason to restart the whole application. A full reload is only an application policy when there is a separate, explicit reason and the user's pending work is protected.

Make recovery observable and testable

Test lifecycle transitions with controlled promises rather than depending on a particular graphics setup. A useful fixture resolves the first adapter or device request only after a second initialization has begun. The expected result is that the first generation never publishes its result; if it created a device, that device is destroyed. A second fixture resolves an older device's lost promise after the replacement is ready; the replacement view and status must remain unchanged.

Also cover the ordinary branches: API absent, adapter request returning null, adapter or device request rejection, loss after successful initialization, user cancellation, retry success, and retry exhaustion. For every branch, assert that page controls remain usable, selected values survive, the status is accessible, and the rendered view agrees with application state. Test reduced-motion behavior separately if recovery uses animation; the recovery itself must not depend on motion to communicate a state change.

Add a replacement fixture with generation A already attached and ready. Start generation B and assert that A's render loop stops, its surface detaches, its owned resources are disposed, and its device is destroyed before B requests an adapter. Run the fixture twice: make B's adapter or device request reject and expect the fallback to remain visible with no A reference; then let B succeed and expect only B to attach. In both cases, resolve A's old lost promise afterward and assert that it cannot change the current view or status.

For a race test, use deferred promises rather than arbitrary timeouts. Start generation A and hold its adapter request. Start generation B, allow B to complete and become visible, then resolve A. The final UI must still show B. Repeat with the first requestDevice() promise held until after B is active; when A finally returns a device, verify that it is disposed and never attached to the renderer. Finally, settle A's lost promise and verify it still cannot change B's state. These checks exercise the stale-work boundary instead of hoping a timing-sensitive test happens to reproduce it.

When creating a test double, expose the same asynchronous boundaries as the API: a deferred adapter request, a deferred device request, and a controllable lost promise for each returned device. Track each call to destroy() and each call that attaches a device to the renderer. The fixture passes only if the obsolete device is destroyed as expected, the current device remains attached, and fallback/status callbacks from the old generation are not called after the replacement is ready. Keep this fixture independent from a real adapter so it can run consistently in automated tests.

Test stop and retry as distinct actions. Stopping should increment the generation, remove the old device, and return to a stable fallback without scheduling a request. A user retry should increment again and enter a visible loading state. If it fails, return to the fallback and leave retry disabled or available according to the product's bounded policy. This makes the expected event sequence explicit: ready, lost, fallback, deliberate retry, then either ready or fallback. A retry count should describe application attempts, not the number of loss events reported by one device.

Verify resource ownership when a device is superseded. A stale device produced by a late promise belongs to the obsolete generation and should be destroyed or otherwise released according to the API lifecycle. The current device must not be cleared by a stale catch block or a stale loss callback. Keep references local to the generation where possible, and check the token before assigning shared state. These are correctness checks: an old rendering loop should not continue drawing into a resource after a new generation has become active.

Rebuilding is not the same as reconnecting an old device. Buffers, textures, pipelines, bind groups, and command encoders belong to the device that created them. After a new device is ready, recreate only the resources required for the active mode and rebuild bindings from application-owned values. If a canvas presentation context uses the device, configure the rendering path for the replacement before submitting new work. Keep CPU-side model data or other authoritative inputs available so the new resources represent the same user-visible task. Do not retain old device resources in a global cache and assume they can be submitted to a new device.

Keep a clear boundary between recovery orchestration and rendering details. A renderer can expose a small operation that prepares resources for a supplied device, while the page decides whether that operation belongs to the current generation. If preparation itself is asynchronous, pass the same generation token through it and check before publishing the prepared renderer. A second loss or a user cancellation can occur while shaders, textures, or other application assets are loading. Those requests should be aborted when possible; otherwise their eventual results must be ignored and released when they belong to a stale generation.

The fallback can preserve meaning without reproducing the exact visual. For a map, it might retain a list of selected places and navigation actions. For a data view, it might use a table with the same filtered values. For an editor, it can preserve the current model and show a non-animated preview. Identify the essential task first, then ensure the alternative represents that task rather than merely stating that graphics failed. Keep controls and instructions in document content so keyboard and assistive-technology users receive the same recovery path.

Make status messages concise and factual. During initialization, tell the user that the enhanced feature is starting. After loss, explain that it stopped and that the rest of the page is available. When a retry is offered, state that it is an attempt, not a guarantee. Avoid showing raw exception strings or browser implementation details to end users; retain only an appropriately minimized operational record if the support process needs one.

Review the recovery contract whenever the graphics feature changes its essential work. A new mode might introduce work that cannot be reconstructed from saved application state; a new export path might make the fallback more valuable than the scene itself. Update the state transition, the data-preservation promise, and the failure tests together. A change to visual quality should not silently turn a previously optional device into a requirement for completing the user's task.

Record expected behavior by transition rather than by machine. A successful initialization means that the requested feature is ready for the current page state. A loss means that the affected graphics work is no longer usable. A successful recovery means that a newly initialized device has rebuilt the required resources from current application data. These statements are stable enough to test across supported environments without inventing a hardware category or a universal recovery guarantee.

Collect only operational information needed to understand the application transition, such as a coarse stage and whether the fallback was shown. Do not gather adapter details, use hardware thresholds, probe the GPU, or treat loss as a fingerprint signal. The public WebGPU specification and MDN GPUDevice.lost reference define the API behavior discussed here; neither promises a specific recovery outcome across browsers or hosts.

Keep the fallback useful

The fallback is part of the feature, not an error page. Keep essential information and actions in ordinary HTML, preserve current selections, and state plainly that the enhanced view stopped while the rest of the page remains available. If some result existed only in GPU memory and cannot be reconstructed, explain that limitation instead of implying that it was saved.

BotBrowser's public WebGPU documentation describes selectable WebGPU modes that may be part of an application's supported test setup. That does not guarantee recovery from device loss, control host resource reclamation, or make the API available in every runtime. Validate the actual user-visible recovery behavior in the environments you support, and keep this application-level contract separate from adapter identity or fingerprinting work.

Sources

#WebGPU#Device Loss#Recovery#Resilient Rendering

Take BotBrowser from research to production

The guides cover the model first, then move into cross-platform validation, isolated contexts, and scale-ready browser deployment.