Skip to main content
Every instance is a Linux machine with a browser on it. The stock agent37-hermes template drives that browser headless: the agent opens pages, clicks, fills forms, and reads what comes back, with nothing for you to watch. Build the desktop recipe instead and the same browser runs on a screen you can open in a tab, watch live, and take control of when a site needs a human.
Paste this into your coding agent

Browsing with no screen

Nothing to configure: agent37-hermes ships Chromium and a browser tool, so “book me a table” or “pull the pricing off these five sites” works from the first chat turn. See Chat. The agent also has the rest of the machine. It can install software with apt-get, run scripts, and serve its own ports, as itself or through exec as root:
curl
What a headless browser can’t do is show you what happened, or let you finish a step it can’t: a sign-in, a code texted to your phone, a CAPTCHA. That is what the desktop is for.

Add a desktop

hermes-vnc-desktop is the stock Hermes image plus a view: a visible Chromium that the agent’s browser tool drives, and noVNC serving that screen on port 6901. Everything the stock template does still happens, because the recipe wraps the stock entrypoint rather than replacing it: the managed model, app connections, web search, and the agent37 CLI the agent uses to schedule itself. Build it into a workspace template once. The build runs on Agent37, so you don’t need Docker:
Then create instances from the name, exactly as you would from a system template:
curl
--default-port 3737 makes the create wait for the gateway, so the instance is ready to chat when it returns. The desktop adds nothing to the bill: a running instance is priced by its shape, not by what runs inside.
Tell the agent it has an audience. Add a line to its instructions saying the user can see its screen, so it browses in the visible browser and asks the user to take over on a sign-in instead of asking for a password in chat.

Open the desktop

Mint a signed URL for port 6901 and change the path from / to /vnc.html, keeping the token:
curl
Open that in a top-level tab. Add &view_only=1 to watch without controlling. Ask the agent to browse something and watch it work.

Embed it in your own app

Load the noVNC client in your own page and connect its WebSocket straight to the instance, with the signed token in the query string. That connection needs no cookie, so it works from your origin in every browser. Your server mints the token for the signed-in user’s own instance and hands the browser only a WebSocket URL:
node
noVNC is plain ES modules, so the page imports a pinned release from a CDN, with nothing to install and no build step. Start in view-only mode. Take over cancels the turn in flight, so you and the agent never fight over the mouse, and turns view-only off; Return control turns it back on:
You and the agent share one browser, so the page you leave open is the one it sees next. Tell it so: send a line of context with the first message after a takeover, saying the user used the computer and it should look at the browser before carrying on.
Don’t iframe the signed URL itself. Its auth rides a SameSite=Lax cookie, which a cross-site frame does not send, so noVNC’s sub-resources come back 401. Connecting noVNC to the WebSocket, as above, needs no cookie. If you would rather frame a page, serve the instance from a custom domain on your own registrable domain, or reverse-proxy port 6901 through your own server with the X-Agent37-Key header attached.

About the token

  • It grants full control. viewOnly is a setting in your page, not a permission: anyone holding the token can connect a VNC client that clicks and types. Mint it only for the instance’s owner.
  • It cannot be revoked. Keep ttl_seconds at 60, the minimum. The token only has to be valid when the socket opens, an open socket keeps working after it expires, and every reconnect mints a fresh one.
  • Opening it wakes a sleeping instance. The edge holds the connection while the instance restores, and the screen comes back as the agent left it.

Worth knowing

  • The login survives. Chromium keeps a persistent profile on the instance’s home volume, so a sign-in you finish during a takeover stays signed in across restarts, updates, and sleep, and the agent’s browser uses the same profile.
  • Auto-sleep works. The desktop and the visible Chromium survive the checkpoint and restore, with the page the agent left open. Connect the view only while it is on screen: it streams even when nothing changes, and that traffic counts as activity, so an open view keeps the instance awake.
  • Give it room. The default 2 vCPU / 4 GB works; use 4 / 8 if the agent opens heavy pages.
  • Ports. 6901 serves noVNC. VNC (5900) and DevTools (9222) stay on loopback inside the instance.
  • Screen size is 1440x900. Add ENV AGENT37_SCREEN_GEOMETRY=1920x1080x24 to the Dockerfile for another size.
  • Crons work as on agent37-hermes, with one difference: on a workspace template, set "agent": "hermes" when you create one, or the run records its session_id only once the turn finishes. See Crons.
  • Telegram webhook ports are wired at create on agent37-hermes only, so a Telegram bot on this template polls and needs auto_sleep off. See Public ports.
A full app built on this is Build your own Dots: a personal agent with its computer beside the chat.