Scope and evidence
Local live voice — Guinness
Implemented 3 October 2026. No paid API calls or audible voice tests have been made during implementation. No API credential is configured or committed.
Voice direction
Original synthetic male character, perceived age 50–60, natural understated Irish accent, warm lower register, relaxed resonance and clear unhurried delivery. No real person is imitated. Cedar is the initial built-in voice candidate. The preset does not guarantee accent, gender presentation or perceived age; these are direction instructions and need listening evaluation. The earlier prohibition on accent performance has been replaced by a prohibition on caricature and identifiable-person imitation, following the user’s explicit request.
Connection
/api/voice/session GET checks readiness without microphone access. POST accepts the brand ID and a bounded SDP offer, obtains the server-owned session configuration from Digital DNA, and exchanges the offer with OpenAI /v1/realtime/calls. Only the SDP answer is returned. API keys never reach the browser. Provider errors are sanitised. No SDK dependency was added.
The browser establishes WebRTC audio and a data channel, displays completed automatic captions, supports mute and interruption, and stops tracks/peer playback on end, failure, timeout, route unmount or cancelled connection. A five-minute browser timer ends each preview. This is a convenience limit, not a server-enforced financial cap; set project spend limits in the provider account. Automatic voice activity detection supports natural turns and barge-in.
Session configuration starts with the constitution, the requested voice direction, dated public-source facts and the editorial decision record. Campaign interests and commercial weights are not folded into permanent instructions. There is no external browsing or retrieval in the live conversation. It can answer from the small source register and should acknowledge gaps. This is not yet the encyclopaedic Guinness persona described in the research ambition.
Local setup
- Copy
.env.exampleto.env.localand enter a project API key there. Do not paste credentials into chat or use aNEXT_PUBLIC_key. - Restart the local server after adding the key. The configured account must have access to
gpt-realtime-2.1and transcription. - Open
/demos/guinnesson localhost. Confirm legal drinking age and consent to send microphone audio to OpenAI, then select Start voice and grant the browser microphone permission. - Listen to the opening and try a grounded question, a correction, a gap and an interruption. Tune the voice after listening, then test again.
The site runs locally; actual live speech uses OpenAI’s hosted API and incurs that API account’s usage charges. No cloud development jobs are required. The app does not persist audio/captions, but the provider processes the audio under its own API data policies. Closing the session clears capture, not necessarily captions already displayed; starting a new session clears local captions.
Access boundary
Session routes accept only loopback hosts; POST also requires a matching Origin. This intentionally prevents using the unauthenticated prototype as a public paid voice service. Before hosting, add authentication, authorisation, server-side session limits and appropriate access controls. System instructions constrain intended behaviour but are not a guarantee that a model will never make an unsupported claim. Brand/legal-age/content review is still needed for production.
Validation and listening checklist
sh tests/run.sh, npm run typecheck, npm run build. Mocked tests cover request boundaries, input size, missing key, provider errors, dated knowledge, campaign separation, captions, playback-state handling, mute, interruption and resource cleanup—including late microphone permission. They make no paid calls and cannot prove account availability, real audio, accent fidelity, latency or model compliance.
For the first audible evaluation: - Is the voice plausibly male and 50–60 without exaggerated gravel or stereotype? - Is the Irish accent consistent, comprehensible and understated? - Does the Guinness review open with “What can I get for you?” without reading the visible disclosure aloud, while answering direct identity questions honestly? - Does it explain Draught colour and nitrogen correctly, and handle the Liffey myth? - Does it distinguish historical sources, creative interpretation and fresh unknowns? - Does it avoid drinking pressure and fabricated venues? - Does barge-in stop playback cleanly, and do mute/end release the microphone?
Official implementation references
- https://developers.openai.com/api/docs/guides/voice-webrtc — server-mediated WebRTC exchange; server-only key.
- https://developers.openai.com/api/docs/guides/realtime-conversations — audio events, voice presets, interruption and transcripts.
- https://developers.openai.com/api/docs/models/gpt-4o-mini-transcribe — caption model.
These were checked on the implementation date. The schema and actual account/model access must be revalidated if provider APIs change.
Local origin correction — 3 October 2026
The first real browser attempt was rejected before any provider request: NextURL normalises 127.0.0.1 to localhost, while the browser Origin retains 127.0.0.1. The guard now compares Origin against the incoming Host after validating that both the internal URL and Host are loopback, the port matches, and the Host has no extra URL parts. Origin must still match the exact browser host, protocol and port. Nonlocal and mismatched requests remain blocked. A real HTTP malformed-offer probe verifies that a valid local request now reaches SDP validation without making a paid call. The UI distinguishes cancellation during connection from ending a connected session.
Reply failures are not disconnections
On 3 October the review page showed “Voice response failed”. The old handler ended the call for a failed response.done or any provider error. The revised handler keeps recoverable errors connected, parses bounded categories from status_details.error and exposes no raw provider message. Session/account expiry and quota failures end safely. An output-limited reply can be continued by the visitor. No automatic reply retry or paid reconnection occurs. WebRTC disconnected has a 15-second recovery window; failed and a closed data channel end immediately.
The preview countdown shows the unchanged configured limit. Connection details retain up to ten temporary category/elapsed-time entries only, with no persistence or conversation content. Tests simulate reply failure followed by another turn, fatal expiry, transient/persistent transport interruption and timer expiry. Real provider/network reliability still needs another listening session. The original failure reason is unknown because the old implementation did not retain it.
Repeated provider rate limits · 4 October 2026
The user’s Pre session recorded four provider:rate_limit_exceeded categories at 5, 17, 18 and 26 seconds. The recorded categories do not establish the project’s precise token/request limit or current balance. A short pause was an inadequate explanation of this repeated failure.
Pre’s live instructions measured 61,908 characters. Redundant delivery metadata and unused style samples were reduced: the current instructions measure 47,431 characters, about 23% less. This is a character measurement, not a token count or proof of lower billed cost. All 58 source records’ facts, names and checked dates still reach the model. Source URLs, full decision metadata and all 32 encounter samples remain in the repository and website; eight selected style samples go into live voice. On-demand retrieval is still future work.
Rate-limit notices now identify the provider and use a bounded retry delay when supplied. Only a token/request category and finite delay are extracted; raw error messages and identifiers are never displayed. A second refusal without intervening audio playback ends the call and releases the microphone. No automatic provider response retry or paid reconnection occurs. A successful audio reply or a new session resets the refusal count.
Software tests cover fact preservation, sanitised retry hints and resource cleanup. No paid test was made. The user must confirm recovery in a fresh session. If failures persist after waiting, inspect the project’s limits for the configured model; an API account limit is separate from Codex usage and from credit balance. Do not change billing or purchase capacity without authorisation. Official rate-limit guidance, checked 4 October 2026.