Troubleshooting Sessions

Most sessions start quickly and run without issues. When something goes wrong, here's how to identify and fix the most common problems.

VM Stuck in "Creating"

Symptom: The VM stays in "creating" status and never reaches "ready."

Cause: Usually a capacity issue -- no servers have enough memory to run your VM right now. Less commonly, a network issue between infrastructure components can prevent the VM from being scheduled.

Fix: Wait a few minutes and retry. If the VM has been stuck in "creating" for more than 5 minutes, destroy it and create a new one. If the problem keeps happening, contact support with the VM or session ID.

Session Waiting for a Machine

Symptom: The session shows Waiting for a machine with an elapsed counter, and messages you send stay Queued.

Cause: Every machine in the fleet was busy when the session asked for one. The session's status is waiting, and waiting_since records when the wait started.

Fix: Usually nothing. The platform retries on its own, the session starts as soon as capacity frees up, and queued messages deliver in order. If you would rather not wait, close the session and start another attempt later.

If capacity never frees up, the platform gives up rather than waiting forever: the session closes with close_reason: resource_exhausted after retrying for over an hour, or waiting_timeout when a wait outlived any retry schedule and was swept. Both carry an error_message saying what happened and send you a notification, and both are resumable, so sending another message starts a fresh attempt.

Slow Boot

Symptom: The session takes noticeably longer to start than usual.

Cause: No snapshot is available for the environment. Without a snapshot, the VM builds from scratch -- cloning your repo, installing dependencies, and starting services. That takes significantly longer than a snapshot boot. A single slow start is also normal right after a snapshot is rebuilt to stay fresh, or is self-healing from one that couldn't be restored; the fresh snapshot swaps back in once the rebuild finishes.

Fix: If it's a one-off right after a rebuild, no action is needed -- the next session boots fast again. Otherwise, make sure your environment has a snapshot in ready status. Check the environment detail page, or inspect .data.base_snapshot.status with GET /v1/accounts/{account}/environments/{environment}. If the snapshot failed or does not exist, trigger a regeneration from the environment settings.

Session Not Responding

Symptom: You send a message but Claude doesn't respond.

Cause: A few things could be happening:

  1. Claude is still working on a previous message. Only one message processes at a time.
  2. The agent crashed mid-turn inside the VM. A failure on the startup path, before the agent begins working, no longer looks like this: the session ends right away with the reason in the thread.
  3. The VM ran out of memory, especially with large builds or many running services.

Fix: Check if the session shows an activity indicator -- if so, Claude is still processing and will respond when done. If the session shows an error instead, read the reason in the thread and see "Session Closed Unexpectedly" below. If there's no activity for more than a minute, try closing and reopening the session. A new VM will boot from the snapshot and your conversation history is preserved.

Messages Not Delivering

Symptom: Messages show as "queued" and never get sent.

Cause: The VM is still booting, or Claude is in the middle of a long response. Messages queue up and deliver once the VM is ready and the current turn finishes. If the current turn was interrupted before it finished, queued messages stay queued rather than being dropped, and deliver when the session resumes.

Fix: Wait for the VM to reach ready status. If the session shows as active and messages remain queued for more than a minute, close the session and start a new one. Your conversation history carries over.

Wallfacer also recovers this on its own. Every message it accepts on an idle or closed session is checked afterwards to confirm the turn actually started; if nothing is running, the platform restarts the session and delivers the message, retrying a few times before giving up. That check runs on a delay of several minutes, so closing and reopening the session is still the faster fix while you're waiting on it.

Agent Doesn't Remember Earlier Turns

Symptom: You reactivate an idle or closed session, and the agent works on your new message with no recall of what it did in earlier turns.

Cause: Reactivating boots a fresh VM and restores the agent's conversation state onto it. When that state doesn't land on the VM, Wallfacer starts the agent on a fresh conversation and delivers your message there rather than failing the turn.

Fix: Nothing in the task is lost. The full history stays readable in the session, and the agent still has the repository, the environment, and the task description. Restate the context it needs in your next message, or point it at the files from the earlier turns. If every follow-up in one environment comes back without recall, contact support with the session ID.

Session Closed Unexpectedly

Symptom: The session moved to "closed" without you closing it.

Cause: This happens for one of a few reasons:

  1. Idle timeout -- default 300 seconds of no activity. Idle sessions usually show idle, not closed.
  2. VM error -- something crashed inside the VM.
  3. Infrastructure reclamation -- the server needed to free resources.
  4. Capacity give-up -- no machine ever came free. The session closes with close_reason: resource_exhausted or waiting_timeout, an error_message explaining the wait, and a notification. See Session Waiting for a Machine.
  5. Turn failure -- the agent's turn failed (for example, an invalid model id, a provider error, or the agent crashing mid-turn). The session closes with close_reason: turn_failed and an error_message describing what went wrong, the app shows an error state, and you get a notification -- the session no longer stops silently. Failures that happen while the turn is still starting up close the session the same way, as soon as they happen, rather than leaving it active until the turn hits its time limit.
  6. The agent's daily spend cap -- the agent has spent its daily model budget, so the turn does not start. This is a turn failure like any other, and the error_message names the agent, what it spent, and when the cap resets.

Fix: Check the session close_reason, and error_message when the close was a failure. Sending a new message can reactivate an idle or closed session when the platform can safely boot a fresh VM; resuming a failed session clears the earlier failure. If the close reason points to setup or snapshot failure, fix that underlying issue first. A capacity close (resource_exhausted or waiting_timeout) needs no fix on your side: send another message when you are ready and the session tries again. If it's turn_failed, the error_message tells you why -- for a bad model id or an un-deployed model, start a session with a valid model; for a spent daily budget, wait for the reset the message names or ask your Wallfacer contact to raise the cap.

Session Went Idle With a Close Reason

Symptom: The session shows idle but carries a close_reason of provisioning_failed or harness_lost, and an error message you did not see happen.

Cause: Wallfacer recovered the session automatically. provisioning_failed means the VM never finished starting up; harness_lost means the agent stopped without reporting how the turn ended. In both cases the session is left resumable rather than closed, so the close reason records what happened without ending the work.

Fix: Send another message. The session boots a fresh VM, the transcript is preserved, and the turn runs again. Read close_reason together with status: a non-null close reason does not by itself mean the session is finished.

Snapshot Boot Failures

Symptom: A session takes a slow, from-scratch boot even though the environment has a snapshot, or the environment briefly shows its snapshot rebuilding on its own.

Cause: The snapshot couldn't be restored -- it may be incompatible with a recent infrastructure update, or tied to a manifest that no longer boots cleanly. Wallfacer no longer fails the session when this happens: it boots from a fresh build and rebuilds the snapshot in the background automatically. The one slow start is that fallback in action.

Fix: Usually nothing -- let the background rebuild finish, and later sessions boot fast again. If the slow, self-healing boot keeps happening, the rebuild itself is failing: regenerate the snapshot from the environment settings page, and if that also fails, check that your manifest's setup commands and services complete successfully, since a snapshot can only be captured once the full environment lifecycle passes.

Large Repository Slow Clone

Symptom: Fresh boots take a very long time (minutes) because of a large repo.

Fix: Make sure you have a snapshot. Snapshots include the prepared checkout, so subsequent boots avoid the full clone and setup path. If you're working with a very large monorepo, consider specifying a branch in your manifest sources or using a shallow clone to reduce the initial clone time.

Setup Command Failure

Symptom: The VM boots but setup commands fail, leaving the environment in a broken state.

Cause: A setup command (like npm install, pip install, or a build script) returned a non-zero exit code. The environment can't reach "ready" status if any setup step fails.

Fix: Check the setup command output in the session activity log. Common causes include:

  • Missing system dependencies your project needs
  • Network issues during package installation
  • Incompatible Node.js, Python, or other runtime versions
  • Package lockfile conflicts

Update the failing command in your manifest's setup section to fix the issue, then regenerate the snapshot.