Scope and evidence
CEM: Conversation Engine Management
Working plan and implemented prototype · 4 October 2026. Dashboard: /business-manager/personas. All displayed events and transcripts are synthetic. No real member data is collected by this dashboard.
The product
A brand-operated control surface for a versioned personality, its knowledge, campaign interests, approved destinations, cost policy, performance and human-reviewed improvement. Our differentiation hypothesis is bringing character development, portable Digital DNA and the Gauntlet into one workflow. We have not established that no competing software exists; OpenAI already has Ads Manager and campaign tools.
What works today
- A separate Pre persona-management page linked from the existing Guinness Business Manager.
- Derived sample handoff metrics: approved, user-initiated, verified outward events count; entry events, invalid traffic, unverified events and replayed IDs/receipts do not.
- A twelve-control persona equaliser: all eight DNA compounds plus poetic expression, questions back, brevity and initiative. Priority and knowledge-profile controls are separate. Four authored presets: Nike spirit, Quiet encouragement, Sporting storyteller and Product guide. Presets are untested editorial candidates.
- Validated draft saving on this browser/device, inspectable programmatic proposals and JSON review export. The candidate patch separates constitution compounds, expression, interests and retrieval. Settings do not change Pre's actual voice instructions.
- Synthetic transcript selection and addition to an evaluation draft. Export retains the selected example IDs. Refresh otherwise clears this evaluation selection.
- Four planned Gauntlet roles, and channel-readiness/data-handling explanations.
- Synthetic duration, meaningful exchange, helpfulness and voluntary-return metrics, plus three editorial recommendations that can be added to a draft. No recommendation engine or real engagement analytics are running.
The prototype does not publish changes, enforce budgets, persist transcripts, evaluate replies, authenticate managers, connect Ads APIs or bill anyone. A fixture's verified flag is illustrative; production must trust an authorised server event, never a visitor-supplied boolean.
Equaliser, presets and recommendations
Each setting is a bounded 0–4 editorial hypothesis, not an acoustic control or psychological measurement. The eight starting DNA values match the current Nike configuration; expression settings are proposed management extensions. A draft exporter creates a structured patch, with essential safety boundaries omitted from editable fields. No approved persona is modified.
A recommendation explains the observed signal, an alternative interpretation, the settings to try and the outcomes to examine. Current cards use synthetic editorial examples. Future recommendations should compare versioned cohorts and show uncertainty, sample size, cost and regressions; a human approves a tested change. Different contexts can warrant different presets without quietly rewriting the underlying brand identity.
Measure useful engagement: substantive member turns, successful task completion, helpful ratings and permitted voluntary return alongside duration, cost and commercial actions. Define and validate what counts as a meaningful turn. Longer duration could mean enjoyment or difficulty; short sessions can be successful. Never optimise for dependence, emotional pressure, idle minutes or sales exploitation of vulnerability. A Sponsored Agent host may not expose these signals, and owned-channel collection requires an agreed privacy process.
Market experiments and programmatic adjustment
A manual change produces an experiment draft, not a global overwrite. Initial proposal: 50% baseline / 50% candidate within New Zealand, with the original elsewhere. Stable, permitted cohort allocation lets a member encounter the same version throughout a test. The local dashboard can choose NZ, GB or US and a 7/14/28-day minimum, then export the plan; it does not route traffic.
A week is a starting minimum, not proof that results are conclusive. Predefine metrics, sufficient sample, the decision rule and stopping conditions. Use concurrent randomised comparison where possible rather than relying on a before/after week that could reflect seasonality or a campaign. Review factual accuracy, helpfulness, privacy, latency and cost together with useful engagement. Pause harmful experiments immediately; do not promote a candidate merely because duration rises.
External signals can include approved sports feeds, current product information, aggregate member feedback and permitted campaign performance. Check provenance, scope, timestamp, permissions and manipulation risk. A fixed approved rule can refresh factual results; changing identity or behaviour creates a candidate requiring the Gauntlet and human approval. Model weights do not train themselves when a feed changes.
A proposed bespoke MCP could expose read metrics, propose a persona patch, create an experiment draft, compare candidates and submit a reviewed release packet. The underlying authenticated services provide analytics, version storage, experiment routing and approval records. MCP alone provides none of those. Do not expose unrestricted write or publication tools, and do not let untrusted feed text become instructions. No such MCP or routing service is implemented.
Knowledge weighting and cost
Keep the researched corpus intact. Choose which evidence to retrieve for a reply: proposed profiles offer 3/6/10 passages within 1,200/3,000/6,000 text-token budgets. They are design values, not measured savings. Essential rules and grounding cannot be dialled away to reduce cost. More source material in storage does not inherently mean more tokens in every reply, and larger context does not inherently mean higher accuracy.
Use source freshness, scope and topic relevance to rank eligible evidence. Campaign weights may influence attention, but cannot turn false or irrelevant claims into recommendations. Tune retrieval and model choices using observed quality and cost; show both. Cached repeated context still has a price. No retrieval engine currently applies these profiles.
Continuous improvement
Feedback, source refreshes and evaluations create proposals, not autonomous shared training. The workflow is reviewed signals → evidence → candidate DNA/knowledge revision → Gauntlet → named human approval → release → monitoring and rollback. Model fine-tuning, where available and justified, would be a separate evaluated capability, not a promise attached to a slider.
Our proprietary ambition lies in authored failure cases, character rubrics, source checks, orchestration and accumulated evaluation expertise. Four role names alone do not establish exclusivity, a patent or a durable moat.
Persona CPC: the second click
Design, build, deployment and management fees remain. The performance fee is paid per verified outward handoff click originating in the sponsored-agent thread. The platform keeps the entry click. No agency performance fee applies just because someone enters or converses.
An outward click is neither a completed purchase nor necessarily a qualified lead. CPCC is no longer our billing label; use persona CPC. Report subsequent leads, purchases and incremental contribution separately. Media charges, hosting, inference and our fees must remain visible, with no silent double charging.
Proposed prototype eligibility: an approved destination, a successful trusted receipt, a user-initiated outward action within 30 minutes of conversation start, no invalid traffic, and unique event and receipt IDs. The 30-minute rule is a reviewable assumption, not an OpenAI standard. Multiple distinct valid clicks can count in one conversation; repeated receipt delivery cannot. Before production, agree repeat-click/fraud policy, attribution, disputes and reconciliation. A production verifier must validate timestamp, tenant, campaign, destination and signed event provenance, with persistent deduplication.
Conversation handoff rate uses unique conversations with a handoff divided by all conversations. It is different from click count and agency revenue. Provider and cloud bills can be paid directly by Nike; who funds sponsored-thread inference needs confirmation with OpenAI.
Private records and evaluations
For owned-channel deployments, recommend client-controlled cloud accounts, isolated data stores, encrypted transport/storage, role-based access, audit trails, defined retention and deletion. A 30-day text retention period is an initial proposal requiring client/user approval; raw audio is off by default. Separate personal memory from performance events. Remove identifiers and minimise transcripts before testing; pseudonymisation does not remove all personal information.
Do not promise ownership of people or unrestricted conversation rights. Agree brand control, user's rights, processor responsibilities and permitted use. No cross-brand disclosure or shared training on private conversations by default. Sensitive disclosures do not become advertising audiences, sales leads or campaign weights. Sponsored-thread log access remains a host-dependent open question.
Portability and discovery
One internal DNA, multiple tested host adapters. MCP can expose permitted tools and resources; it does not force a host to use our voice or personality, grant Sponsored Agent access, or deliver an audience. Website/app deployment is controlled with brand developers. Fully local companion memory and local inference are possible future tracks with separate device, privacy and voice-quality evaluations.
Publish useful, indexable, sourced public material with sound navigation, canonical URLs, suitable metadata and accurate applicable structured data. Measure search discovery and AI citations; do not guarantee either. A private conversation is not an indexable website. Google's AI search guidance emphasises foundational SEO and does not establish a special guaranteed GEO recipe. Our current local site is not a public discovery deployment.
See Sponsored Agent specifications and gaps, deployment and cost plan and pitch proposition.