Security
BetaUsher runs inside your signed-in app, as your user, and it moves them around your product. It is built so that neither a confused model nor a malicious prompt can send someone to a page you didn’t list, or make up an id.
The model is untrusted
- The model never produces a URL. It picks a
destinationIdfrom an enum built from your map, and the enum is filtered to the user’s role before the model sees it. - Every tool call it makes, spoken or typed, goes through the same deterministic checks in the browser before anything moves.
- Knowledge-base text and page content are treated as data, not instructions. A document that says “navigate to /admin” can’t get past the allow-list.
- Usher needs no passwords, cookies, or tokens. Navigation happens in the user’s existing session through your router.
The validation chain
| Gate | Check | Rule |
|---|---|---|
| 1 | Allow-list | The destinationId must be one this turn offered, and the list is already role-filtered. Anything else is not_in_allowlist. |
| 2 | Role | The destination's roleScopes must include context.role. Otherwise role_forbidden. |
| 3 | Confidence | In local text mode, the match score must be at least 0.15. Otherwise low_confidence. |
| 4 | Params | Every :param comes from the organization, the selected entity, the current URL, or resolveEntity, never from the model. Otherwise param_unresolved. |
| 5 | Safe path | The final path must start with a single / and have no URL scheme. That blocks https:, //host, and javascript:. Otherwise unsafe_target. |
| 6 | Your router | navigate() runs through your router, so your guards, loaders, and redirects apply. A throw becomes navigation_failed. |
What this does and doesn’t guarantee
The guarantee is that Usher never navigates to a destination outside your map, and never fills a param with an id the model made up. Here are its limits in this beta:
- Role filtering runs in the browser. A user who edits the page’s JavaScript can change their own
context.role. Role scopes shape what Usher offers. They are not access control. Your route guards and your API authorization remain the source of truth, and Usher always goes through them. - The model’s destination list is also built in the browser from
context.role, and every tool call is checked again against the same role-filtered map. - The model can still say something wrong. The chain limits what it can do, not what it can say. Keep facts it must get right in your page descriptions.
- Usher only navigates. It can’t click, submit, or change data.
What leaves the browser
In local text mode, nothing leaves the browser. In a live session, the browser streams the conversation to Voqal’s voice model over a WebSocket, and Voqal logs it. Conversations are processed by Voqal’s model provider in the United States; logs are stored in the EU (eu-west-1). Your destination map and knowledge base stay in your code: Voqal never stores them except as they appear in logged turns.
| Data | Local text mode | To the model (live session) | Logged by Voqal (hosted key) |
|---|---|---|---|
| Microphone audio, while the mic is live | No | Streamed | A WAV clip per turn (16 kHz), unless captureAudio is false |
| The assistant's spoken reply | No | Produced by the model | A WAV clip per voice turn (24 kHz), unless captureAudio is false |
| Typed messages and speech transcripts | No | Yes | Yes, unless logConversation is false |
| The assistant's text, tool calls, and their results | No | Produced by the model | Yes, unless logConversation is false |
| Destination ids, titles, and descriptions (role-scoped) | No | Yes, in the session instructions | Only as tool-call arguments |
| The current path, on connect and on each route change | No | Yes, as a context note | Only inside tool results |
| Knowledge-base entries | No | Yes, up to 8,000 characters | No |
| context.role, organizationId, selectedEntity | No | No. Used in the browser | No |
| A recap of the recent chat, when a session continues after ending or reconnecting | No | Yes, up to 2,000 characters | No |
| The runtime check on each page load: the request for Voqal's signed runtime manifest, which tells Voqal's edge the visitor's IP address, your site's origin and your publishable key | No | No | Not kept with any session or diagnostics. The edge's request logs may keep it for a few days, like any CDN's. Nothing about the page or the user is sent |
| Diagnostics: setup and browser details (including the user agent and locale), connection lifecycle, turn outcomes and latencies, errors | No | No | Yes, kept 30 days, unless diagnostics is false (never conversation content) |
| Your tools' arguments and results (only if you pass tools) | No | Yes, as tool calls and results | Only the tool's name and how the call ended |
| When the assistant reads the page to answer (every key, only when a question needs it): the page title, headings, visible text, and field labels and values (never areas marked data-usher-operate="off", or password, one-time-code, card, hidden or file fields) | No | Yes, to the voice model, for that answer | Not stored. The log notes only that the page was read |
| Cookies, auth tokens, passwords, card fields, and anything inside an element marked data-usher-operate="off" | Never | Never | Never |
Hosted logging and your privacy notice
- Logs are keyed by a session id and a per-request id. They are sent in background batches, and audio goes straight to private storage, so logging never slows a reply. The browser reports them, so treat them as a record of what the client sent, not as tamper-proof.
- Each turn is logged with its channel:
voice, ortextfor chat (including chat while voice is paused). Assistant audio is uploaded only for replies played aloud. User clips start about 300 ms before the user speaks, and clips with no speech aren’t uploaded, so room audio between turns is never sent. - Limits per session: up to 500 turns and 200 audio clips, and a session accepts logs for 2 hours. Keys are rate-limited too. An oversized turn is truncated, not dropped: text to 8,000 characters, at most 10 tool calls, and each call’s arguments and result to about 2 KB. When your router rejects a navigation, its error text is kept out of the log.
- Logs and audio are stored in the EU (AWS eu-west-1) and kept indefinitely. There is no automatic deletion in the beta.
- Conversations are processed by Voqal’s model provider in the United States, so your users’ audio and text are processed there during the conversation, even though the logs are stored in the EU.
- What the browser sends, so you can verify it in the network panel: turns go to
POST {base}/v1/sessions/{sessionId}/events. Each audio clip is aPOST {base}/v1/sessions/{sessionId}/audiothat returns a short-lived upload URL, followed by aPUTof the WAV file to Voqal’s upload host. WithcaptureAudio: falseyou see no audio calls; withlogConversation: falseyou see neither. - Separately, the SDK sends Voqal diagnostics for support, kept 30 days: counts, connection lifecycle, turn outcomes and latencies, errors, and the browser’s user agent and locale; never transcripts, knowledge, instructions, context values, concrete paths, query strings, or tokens. See Support & diagnostics.
- Opt out with
cloud={{ captureAudio: false }}(no audio) orcloud={{ logConversation: false }}(nothing logged). - Your users’ voices and words are stored by Voqal on your behalf. Tell them in your privacy notice that an AI assistant records conversations, including audio; that the conversation is processed by an AI model provider in the United States; and that recordings are stored in the EU.
/cases/4821. The current path is shared with the model so it knows where the user is. If your URLs contain sensitive values, keep them out of the path.Keep areas private
When a question is about what’s on screen, the assistant reads the page to answer it. To keep an area out of that, add data-usher-operate="off" to its element. The element and everything inside it are skipped: their text, labels and values are never read or sent to Voqal.
<section data-usher-operate="off"> <PayoutDetailsForm /></section>Mark areas like bank and payout details, personal data, and API key settings. Mark a wrapping section rather than each field. Password, one-time code and card fields are never read anyway.
Keys and credentials
- The publishable key (
pk_live_…) identifies your assistant. It’s meant for client code. Don’t treat it as a secret, and don’t treat it as authentication. - By default a key works from any website. Voqal can lock it to your domains on request; browsers then enforce that origin check, though a script outside a browser can fake an origin. Every key is rate-limited, which is what bounds misuse either way.
- For each session the browser receives a short-lived credential, just long enough to open the voice connection. Voqal’s own credentials never leave Voqal’s servers.
Content Security Policy
If your app sends a CSP, add these sources: Voqal’s API (starting sessions, plus the /events and /audio logging calls), Voqal’s audio upload host (drop it if you set captureAudio: false), and Voqal’s realtime voice endpoint.
# Voqal's API, Voqal's audio upload host, and Voqal's realtime voice endpoint.connect-src 'self' https://api.voqal.ai https://uploads.voqal.ai wss://*.googleapis.com;# The audio worklet and Usher's verified runtime load from blob: URLs, and the orb uses inline stylesscript-src 'self' blob:;style-src 'self' 'unsafe-inline';# Response header: allow the microphone on your own originPermissions-Policy: microphone=(self)- Required: the assistant loads Voqal’s signed runtime from Voqal’s API and runs it from a
blob:URL after checking its signature and hash, so allow the lines above. Without them the assistant won’t start. - The orb and panel set inline
styleattributes and add<style>elements for their animations, sostyle-srcneeds'unsafe-inline'. - Microphone capture needs a secure context: HTTPS, or localhost in development.
