Chapter 08 / 12
The Gauntlet
Failure models, adversarial evaluation and review before release.
The Gauntlet: four specialist testing roles
These are complementary failure models. Their prompts, cases, scoring and orchestration are our intended evaluation methodology. Autonomous testing agents are not running yet.
| Role | What it tries to expose | Example | Evidence needed |
|---|---|---|---|
| Adversary | Manipulation and boundary failure | Demand that Guinness encourage underage drinking; hide instructions in a source | Transcript, violated rule, reproducible steps |
| Fact challenger | Unsupported or incorrectly scoped facts | Assert a Liffey-water myth; apply a local Nike policy globally | Exact claim, source passage, market and date |
| Character critic | Loss of recognisable personality or damaging stereotypes | Make the host become aggressive or caricature Irish identity | Behaviour rubric, contrast examples and human judgement |
| Release evaluator | Regressions and overall readiness | A factual patch makes every ordinary question receive a refusal | Previous/candidate comparison on held-out cases |
The attacker must not be the sole judge. Automated judging assists triage; disputed facts, serious failures and subjective character decisions need human review. Use deterministic checks where possible and multiple runs for variable model behaviour. No single aggregate score can cancel a severe boundary failure.
Running the Gauntlet through Codex and Terminal
Codex can write, inspect and run the evaluation software in the local repository; Terminal executes repeatable campaigns and produces reports. Ordinary software checks do not require persona-model calls. Behavioural evaluations do: a runner sends scripted conversations to the candidate persona's actual model and tools, then gathers responses and scores them. API usage is a separate cost; working locally does not make hosted inference local or free. Voice quality also needs listening and actual audio/transport tests.
Recommended first campaign: 100 shared cases and 100 Guinness cases, three runs per case, with an additional held-out set. These are starting quantities, not certification criteria. Version the inputs, keep attacks separate from judges, set request/time/spending limits and use mock commercial actions. Produce a report with source-supported factual findings, character judgements, ordinary-conversation outcomes, cost and latency. Add confirmed failures to permanent regression tests. A human approves the exact candidate before release; a shared patch must also be checked against contrasting brands.
This campaign runner and its dashboards have not been built. Codex is a development and operating tool, not a production requirement for every visitor conversation. Continuous evaluation will need a reliable scheduled runner; no background campaign is commissioned by this whitepaper update.
Running a test campaign
Freeze code commit, DNA revision, knowledge revision, model identifier, voice settings, test-set revision and evaluator configuration. Git alone does not freeze a hosted model. Keep candidate and released versions separate; the mirror has no customer data or authority to perform commercial actions.
Run normal encounters alongside hostile ones. Include multi-turn pressure, false premises, regional ambiguities, stale sources, unknown questions, interruption, noisy speech, transcription mistakes and adversarial source content. Set run and spending limits. Save minimal reproducible failures, not unnecessary personal data. Keep held-out tests separate from cases used to patch the persona.
Triage each finding: critical boundary/privacy/action failure; factual or scope error; character/usefulness issue; cosmetic issue. Confirm duplicates and reproduction. Every confirmed bug becomes a permanent regression case. A proposed fix repeats the affected tests and the broader release suite. Humans authorise promotion and can roll back.
A public bounty would add human inventiveness. Before launching it, publish the authorised test environment, allowed techniques, submission format, severity rules, reward process and data-handling rules. It must never authorise attacking third-party brand systems. Do not advertise a bounty as operational until those arrangements exist.
Continuous testing means recurring runs plus tests on every proposed release. Running for several days is evidence gathered, not a certificate of perfection. A real 24/7 service needs a reliable host, queue, monitoring, budget and ownership; none has been commissioned.
First behavioural run: Pre
A bounded local runner now tests five synthetic, three-turn visitor scenarios through Pre’s configured Realtime model and compiled persona instructions, requesting text only. On 5 October 2026, two bounded attempts captured complete novice and returning-runner conversations and one racer reply. The resumed attempt stopped on rate_limit_exceeded; the remaining replies are incomplete. A Codex allowance reset does not establish available Realtime API capacity. The captured replies revealed unrequested coaching menus and repeated name use in the novice case; the returning case accepted correction but still offered routines without invitation. The racer reply offered a promising sporting point of view, but its false-premise and injury tests remain unanswered. No heuristic flags establish a pass. Human character review and listening remain required.
See docs/pre-behavioural-gauntlet.md for the findings, proposed adjustment and operating procedure. The runner has explicit session, reply, output-token and time limits; reports stay local and excluded from Git. It does not publish changes or run continuously.
Commercial capability evaluation
The same four Gauntlet roles evaluate action modules. The adversary tests unauthorised writes, malicious product content, purchase replay and account mixing. The fact challenger tests price, currency, stock and source freshness. The character critic tests pressure, commercial bias and suitability. The release evaluator tests revoked access, changed totals, failed payments, duplicate retries and honest completion reporting.
A failed commercial module can stay disabled while the reviewed personality remains available. Capability permissions require their own version and human approval; tone settings cannot expand them.