Security

Beta
Last updated  Oct 5, 2026

Usher runs inside your signed-in app, as your user, and it moves them around your product. It is built so that neither a confused model nor a malicious prompt can send someone to a page you didn’t list, or make up an id.

The model is untrusted

  • The model never produces a URL. It picks a destinationId from an enum built from your map, and the enum is filtered to the user’s role before the model sees it.
  • Every tool call it makes, spoken or typed, goes through the same deterministic checks in the browser before anything moves.
  • Knowledge-base text and page content are treated as data, not instructions. A document that says “navigate to /admin” can’t get past the allow-list.
  • Usher needs no passwords, cookies, or tokens. Navigation happens in the user’s existing session through your router.

The validation chain

GateCheckRule
1Allow-listThe destinationId must be one this turn offered, and the list is already role-filtered. Anything else is not_in_allowlist.
2RoleThe destination's roleScopes must include context.role. Otherwise role_forbidden.
3ConfidenceIn local text mode, the match score must be at least 0.15. Otherwise low_confidence.
4ParamsEvery :param comes from the organization, the selected entity, the current URL, or resolveEntity, never from the model. Otherwise param_unresolved.
5Safe pathThe final path must start with a single / and have no URL scheme. That blocks https:, //host, and javascript:. Otherwise unsafe_target.
6Your routernavigate() runs through your router, so your guards, loaders, and redirects apply. A throw becomes navigation_failed.

What this does and doesn’t guarantee

The guarantee is that Usher never navigates to a destination outside your map, and never fills a param with an id the model made up. Here are its limits in this beta:

  • Role filtering runs in the browser. A user who edits the page’s JavaScript can change their own context.role. Role scopes shape what Usher offers. They are not access control. Your route guards and your API authorization remain the source of truth, and Usher always goes through them.
  • The model’s destination list is also built in the browser from context.role, and every tool call is checked again against the same role-filtered map.
  • The model can still say something wrong. The chain limits what it can do, not what it can say. Keep facts it must get right in your page descriptions.
  • Usher only navigates. It can’t click, submit, or change data.

What leaves the browser

In local text mode, nothing leaves the browser. In a live session, the browser streams the conversation to Voqal’s voice model over a WebSocket, and Voqal logs it. Conversations are processed by Voqal’s model provider in the United States; logs are stored in the EU (eu-west-1). Your destination map and knowledge base stay in your code: Voqal never stores them except as they appear in logged turns.

DataLocal text modeTo the model (live session)Logged by Voqal (hosted key)
Microphone audio, while the mic is liveNoStreamedA WAV clip per turn (16 kHz), unless captureAudio is false
The assistant's spoken replyNoProduced by the modelA WAV clip per voice turn (24 kHz), unless captureAudio is false
Typed messages and speech transcriptsNoYesYes, unless logConversation is false
The assistant's text, tool calls, and their resultsNoProduced by the modelYes, unless logConversation is false
Destination ids, titles, and descriptions (role-scoped)NoYes, in the session instructionsOnly as tool-call arguments
The current path, on connect and on each route changeNoYes, as a context noteOnly inside tool results
Knowledge-base entriesNoYes, up to 8,000 charactersNo
context.role, organizationId, selectedEntityNoNo. Used in the browserNo
A recap of the recent chat, when a session continues after ending or reconnectingNoYes, up to 2,000 charactersNo
The runtime check on each page load: the request for Voqal's signed runtime manifest, which tells Voqal's edge the visitor's IP address, your site's origin and your publishable keyNoNoNot kept with any session or diagnostics. The edge's request logs may keep it for a few days, like any CDN's. Nothing about the page or the user is sent
Diagnostics: setup and browser details (including the user agent and locale), connection lifecycle, turn outcomes and latencies, errorsNoNoYes, kept 30 days, unless diagnostics is false (never conversation content)
Your tools' arguments and results (only if you pass tools)NoYes, as tool calls and resultsOnly the tool's name and how the call ended
When the assistant reads the page to answer (every key, only when a question needs it): the page title, headings, visible text, and field labels and values (never areas marked data-usher-operate="off", or password, one-time-code, card, hidden or file fields)NoYes, to the voice model, for that answerNot stored. The log notes only that the page was read
Cookies, auth tokens, passwords, card fields, and anything inside an element marked data-usher-operate="off"NeverNeverNever

Hosted logging and your privacy notice

  • Logs are keyed by a session id and a per-request id. They are sent in background batches, and audio goes straight to private storage, so logging never slows a reply. The browser reports them, so treat them as a record of what the client sent, not as tamper-proof.
  • Each turn is logged with its channel: voice, or text for chat (including chat while voice is paused). Assistant audio is uploaded only for replies played aloud. User clips start about 300 ms before the user speaks, and clips with no speech aren’t uploaded, so room audio between turns is never sent.
  • Limits per session: up to 500 turns and 200 audio clips, and a session accepts logs for 2 hours. Keys are rate-limited too. An oversized turn is truncated, not dropped: text to 8,000 characters, at most 10 tool calls, and each call’s arguments and result to about 2 KB. When your router rejects a navigation, its error text is kept out of the log.
  • Logs and audio are stored in the EU (AWS eu-west-1) and kept indefinitely. There is no automatic deletion in the beta.
  • Conversations are processed by Voqal’s model provider in the United States, so your users’ audio and text are processed there during the conversation, even though the logs are stored in the EU.
  • What the browser sends, so you can verify it in the network panel: turns go to POST {base}/v1/sessions/{sessionId}/events. Each audio clip is a POST {base}/v1/sessions/{sessionId}/audio that returns a short-lived upload URL, followed by a PUT of the WAV file to Voqal’s upload host. With captureAudio: false you see no audio calls; with logConversation: false you see neither.
  • Separately, the SDK sends Voqal diagnostics for support, kept 30 days: counts, connection lifecycle, turn outcomes and latencies, errors, and the browser’s user agent and locale; never transcripts, knowledge, instructions, context values, concrete paths, query strings, or tokens. See Support & diagnostics.
  • Opt out with cloud={{ captureAudio: false }} (no audio) or cloud={{ logConversation: false }} (nothing logged).
  • Your users’ voices and words are stored by Voqal on your behalf. Tell them in your privacy notice that an AI assistant records conversations, including audio; that the conversation is processed by an AI model provider in the United States; and that recordings are stored in the EU.
Paths can contain ids, like /cases/4821. The current path is shared with the model so it knows where the user is. If your URLs contain sensitive values, keep them out of the path.

Keep areas private

When a question is about what’s on screen, the assistant reads the page to answer it. To keep an area out of that, add data-usher-operate="off" to its element. The element and everything inside it are skipped: their text, labels and values are never read or sent to Voqal.

private-area.tsx
<section data-usher-operate="off">  <PayoutDetailsForm /></section>

Mark areas like bank and payout details, personal data, and API key settings. Mark a wrapping section rather than each field. Password, one-time code and card fields are never read anyway.

Keys and credentials

  • The publishable key (pk_live_…) identifies your assistant. It’s meant for client code. Don’t treat it as a secret, and don’t treat it as authentication.
  • By default a key works from any website. Voqal can lock it to your domains on request; browsers then enforce that origin check, though a script outside a browser can fake an origin. Every key is rate-limited, which is what bounds misuse either way.
  • For each session the browser receives a short-lived credential, just long enough to open the voice connection. Voqal’s own credentials never leave Voqal’s servers.

Content Security Policy

If your app sends a CSP, add these sources: Voqal’s API (starting sessions, plus the /events and /audio logging calls), Voqal’s audio upload host (drop it if you set captureAudio: false), and Voqal’s realtime voice endpoint.

Content-Security-Policy
# Voqal's API, Voqal's audio upload host, and Voqal's realtime voice endpoint.connect-src 'self' https://api.voqal.ai https://uploads.voqal.ai wss://*.googleapis.com;# The audio worklet and Usher's verified runtime load from blob: URLs, and the orb uses inline stylesscript-src 'self' blob:;style-src 'self' 'unsafe-inline';# Response header: allow the microphone on your own originPermissions-Policy: microphone=(self)
  • Required: the assistant loads Voqal’s signed runtime from Voqal’s API and runs it from a blob: URL after checking its signature and hash, so allow the lines above. Without them the assistant won’t start.
  • The orb and panel set inline style attributes and add <style> elements for their animations, so style-src needs 'unsafe-inline'.
  • Microphone capture needs a secure context: HTTPS, or localhost in development.
© 2026 VoqalVoqal SDK & engine documentation