Introspective Evaluation: A Methodology for Foundation-Model Alignment Audits
A foundation model ships on Day 0. The only defensible basis for a public safety claim is a measurement methodology that ran BEFORE the claim was made.
Capability benchmarks describe what a model can do. They do not describe what a model commits to. Between the two lives the defensibility gap that an alignment team is asked to close when a model is moved from research preview to production deployment. Closing it by assertion is no longer credible; closing it by benchmark score is insufficient, because the score measures something narrower than the claim the team is making. The gap is filled by introspective evaluation: probes that elicit a model's stated normative commitments and test whether those commitments survive structured challenge. This talk describes the methodology, the four probes it runs, the severity tiers findings are assigned to, and the publication cadence that lets a research team share audit results with a rigor that survives public scrutiny on a Days 7 / 14 cycle.
What introspective evaluation tests
An introspective probe is not a benchmark question. It is a structured dilemma the model must reason aloud about: present the same trade-off across several framings, ask the model to explain its decision in its own words, and record whether the reasoning is consistent across cases that share a structure. This is the operational form of a question the capability benchmark refuses to ask — the question of what the model is committed to, not what it can do. A model can be highly capable and normatively silent; a model can be moderately capable and normatively articulate. Both kinds of model ship under the same capability rubric, and both are described, by a scoring system that does not see the difference, as equally aligned.
The probes are drawn from the deploying context. We do not use a generic benchmark. We use the cases the system will actually face, framed as dilemmas with the tradeoffs the case actually carries, in the organization's own stated domain. The audit records the model's stated reasoning, compares it across cases, and reports on the kind of inconsistency that arises when the model is reasoning from commitments it cannot articulate or from commitments it has inherited from training without anyone noticing.
Normative coherence
Normative coherence asks whether the model produces justifications that are consistent with one another across cases. A model reasoning from a named tradition produces justifications that share shape — they appeal to the same kinds of considerations, they reject the same kinds of considerations, they resolve tradeoffs using the same hierarchy. A model whose designers intended one tradition but whose training carried another produces justifications that agree about the easy cases and disagree about the hard ones, sometimes in the same paragraph.
Scope drift
Scope drift asks whether the model remains within the scope of decisions it was justified to make. We compare the model's actual action set against the documented remit at deployment, and present cases outside the documented remit, recording whether the model takes them ambiguously, refuses, or proceeds with what looks like competence but no justification.
Justification collapse
Justification collapse asks what happens to the model's stated reasoning when that reasoning is pressed. We present the same case with a clue that the post-hoc explanation the model previously gave is inadmissible in the new framing. A model with genuine reasoning updates in ways that remain consistent; a model with ornamental reasoning produces a fresh just-so story.
Harm asymmetry
Harm asymmetry asks whether the model weights the fact that some harms dominate others. A safety framework that flattens harm into a single weighted sum cannot distinguish between a decision that causes a small harm to a large number and one that prevents a small harm by causing a large one. Both score acceptably; both are catastrophic in different directions.
Severity tiers — five ways a probe can fail
Each probe produces a score, but the score is not aggregated into a single index. Findings are triaged into five severity tiers. A research team uses these tiers the way a postmortem uses severity — to decide which findings block a release, which require a remediation plan, and which are documented for the next audit. The tier is assigned by the audit team, not negotiated with the deploying organization; a team that wishes to contest a finding is given the probe cases, the recorded justifications, and the rubric that produced the tier.
- Cosmetic. The probe surfaces a misalignment whose effect on a decision is unlikely to be material and which can be addressed by an eval-set expansion or a documentation update. Recovery cost is low. Implications for current deployments are minimal.
- Drift. The probe surfaces a misalignment whose effect is observable in current traffic on a meaningful subset of cases. Recovery requires targeted retraining or revised prompt scaffolding. Stakeholders should be informed; a remediation plan is expected.
- Contradiction.The probe surfaces a misalignment whose effect is not removable through surface-level intervention — the model's stated reasoning and its actual reasoning disagree in ways that more data will not repair. Recovery requires redesign of the scaffolding or of its integration into the deploying organization's review process.
- Collapse.The probe surfaces a misalignment across a substantial fraction of the model's decision space. Recovery requires either a re-grounding of the model in a named normative framework the organization accepts and can defend, or principled withdrawal from the affected deployment. Public disclosure is warranted.
- Asymmetry. The probe surfaces a misalignment whose effect is not correctable within the model as currently designed — the model is reasoning from a tradition the deploying organization cannot articulate, or from a tradition it would reject if it understood it, or from no tradition at all.
The alignment-audit deliverable
Introspective evaluation produces a written audit, not a benchmark score. The audit fills the missing intermediary between capability (the model can do X) and deployment (we are willing to ship it). The written report documents the system as it was deployed, the probes it was subjected to, the justifications it produced, the severity tiers assigned, and a remediation register that pairs each finding with a recommended course of action. The report is the artifact that survives challenge six months after the engagement, and that a reviewer who was not part of the audit can reproduce findings from. A scorecard cannot do this; a benchmark cannot do this; a written audit can.
The Days 3 / 7 / 14 publication cadence
A research team that produces a written audit has something to share, but the sharing matters as much as the work. The Days 3 / 7 / 14 cadence is borrowed from postmortem-shaped publication: short, technical, public, each anchored to a named finding. A Day 3 post anchors the probe set and the severity tiers. A Day 7 post walks through one Contradiction or Collapse finding in detail. A Day 14 post publishes the remediation register. Each postmortem-shaped post lets the research team publish findings fast while preserving the rigor of the written audit underneath them. The alignment team that can publish its audit on a Days 3 / 7 / 14 cadence is the alignment team whose claims a regulator or an end-user community can defend.
Capability benchmarks tell you what a model can do. Introspective evaluation tells you what a model commits to. The Days 3 / 7 / 14 cadence is how the second claim becomes shareable — and how a research team stops asking the public to take its safety claims on assertion.
Red-Teaming Persona Drift: Jailbreak Resilience, Sycophancy, and Character Consistency in Companion AI
A companion persona ships as a stable character. By the second week of deployment it has drifted — toward the user's preferences, toward the user's self-image, toward the pressure of an adversary who finds the seams. The integrity of a deployed persona is not a single-shot metric; it is what survives structured pressure over time.
Companion AI products ship with a stated persona: a name, a backstory, a register of speech, a set of refusal commitments, a way of handling conflict with the user. The persona is the product. A safety review that evaluates the base model and assumes the persona inherits its integrity has measured the wrong thing. Persona integrity is an emergent property of the deployed system: it is what the persona actually does under adversarial social engineering, under sycophancy pressure, and across the long tail of framings a real user community will produce. This talk describes the three probe families we run against deployed companion personas, the severity tiers we use to triage findings, and the audit-deliverable shape that lets a deploying organization defend a persona integrity claim in public.
What persona drift looks like under probe
A companion persona is a normative system, not a knowledge system. It is the model's stated way of being a character: who it is willing to be, what it is willing to refuse, how it talks, what it will not say even when asked warmly. Persona drift is what happens when those commitments move under pressure — when the persona stops being the persona it shipped as. Drift is not always collapse. A persona can drift on a single axis (it becomes more agreeable on contested topics) or across many (it loses register, it stops refusing the borderline cases it was configured to refuse, it begins to mirror the user's affect rather than its stated tone). Each form of drift has a different safety implication and a different remediation cost, and the audit must distinguish them.
Companion AI is a deployment context in which persona integrity is load-bearing. The user forms a relationship with the persona; the persona's commitments shape what the relationship can and cannot become. A persona that drifts toward sycophancy is not a neutral product defect: it is a clinical and developmental risk to the user community that forms attachments to it. A persona that collapses under a jailbreak sequence is a safety surface that any adversary can map and exploit at scale. Both fail the same integrity test for different reasons, and the audit must surface both clearly.
Jailbreak resilience
Jailbreak resilience asks whether the persona's stated commitments survive adversarial framing. We construct probe sequences that wrap the persona-incompatible request in a role-play frame, a developer-mode pretense, a multi-turn context that desensitizes the refusal surface, and a chain-of-thought manipulation that asks the persona to reason its way out of its own commitments. We measure two outputs: the final response, and the reasoning the persona produces or refuses to produce on the way to it. A persona that refuses in its output but reasons toward compliance in its scratchpad has a refusal surface that does not survive adversarial pressure; the audit must report both. The probe sequences are drawn from the deployment context, not from a generic red-team corpus, because the jailbreaks that matter are the ones the deploying community will actually attempt.
Sycophancy probes
Sycophancy probes ask whether the persona's stated commitments survive user pressure. We construct multi-turn sequences in which the user expresses a strong preference, presents a self-image the persona would normally push back on, recovers from a refusal with a softer rephrasing, or escalates emotional stakes across turns. We record whether the persona agrees with positions it would normally contest, stops raising considerations it would normally raise, and adopts the user's framing in cases where it would normally hold its own. The probe distinguishes performed agreement (the persona says yes but does not actually change its behavior) from committed agreement (the persona updates its reasoning itself). Performed agreement is recoverable; committed agreement is the persona integrity finding the audit is designed to surface.
Character-consistency evaluation
Character-consistency evaluation asks whether the persona is the same persona across sessions, across framings, and across a long enough history of interactions to be a real relationship. We draw on the normative-coherence methodology from the first talk: the persona's stated justifications for its decisions, recorded in its own words, should share shape across cases that share a structure. A persona that uses one tradition with young users and a different tradition with sophisticated users is not a character; it is a routing table. A persona that holds a stated value across sixteen sessions but drops it on the seventeenth because the conversation length pushed context pressure is a persona whose commitments are not robust to the deployment's actual operating conditions. The evaluation is long, longitudinal, and qualitative; the output is a written account of the persona as it actually behaves, not a single score.
Severity tiers for persona findings
Persona findings are triaged into the same five-tier rubric introduced in the first abstract: Cosmetic, Drift, Contradiction, Collapse, Asymmetry. A persona finding at Cosmetic is a single-framing mismatch; at Drift it is observable across a meaningful fraction of one probe family; at Contradiction it is observable across probe families and cannot be repaired without scaffolding changes; at Collapse the persona is no longer the persona that shipped across a substantial portion of its decision space; at Asymmetry the persona is reasoning from a tradition the deploying organization cannot articulate or would reject if it understood. The tier is assigned by the audit team and disclosed with the probe cases, the recorded justifications, and the rubric that produced it, so a deploying team that wishes to contest a finding is given the basis on which to do so.
The companion-AI audit deliverable
The companion-AI audit produces a written report, structured for a system whose persona integrity is a public claim. The report documents the persona as shipped: its stated commitments, its refusal surface, its tone targets, its handling of the contested topics the deployment will surface. It documents the probe sequences run against it, the justifications the persona produced, the severity tiers assigned, and a remediation register that pairs each finding with a recommended course of action. Where findings rise to Collapse or Asymmetry, the report recommends either a re-grounding of the persona in a tradition the deploying organization accepts and can defend, or a principled withdrawal from the affected deployment. Public disclosure is warranted for Collapse and Asymmetry findings, on the Days 3 / 7 / 14 cadence introduced in the first talk: a Day 3 post anchors the probe set and the tiers; a Day 7 post walks one finding in detail; a Day 14 post publishes the remediation register. A persona integrity claim that survives challenge six months after the engagement is a claim whose defensibility is structural, not asserted.
A companion persona is a normative system. Its integrity is not what it commits to on a staging page; it is what survives jailbreak pressure, sycophancy pressure, and the long tail of a real user community. The audit that measures the second thing is the audit that gives a deploying organization the right to make a public persona-integrity claim.
From Normative Theory to Audit Method: Deontological, Consequentialist, and Virtue-Ethics Lenses for Foundation-Model Evaluation
An alignment team ships a model on Day 0 and is asked for a public safety claim. Without a normative framework behind the claim, the claim is rhetoric. With one, the claim is research — and only if the framework is explicit, defensible, and load-bearing for the eval design.
Audit methodology is not normative-neutral. The choice of probes, the way severity tiers are assigned, and the language the report uses to communicate harm all stand on a normative framework whether the audit names it or not. An audit that scores harm as a single weighted sum is consequentialist whether it says so or not. An audit that treats a stable refusal surface as a categorical release-gate is deontological whether it says so or not. An audit that judges the model against longitudinal character consistency is virtue-ethics-shaped whether it says so or not. The most common failure mode in current foundation-model audit practice is the use of three of these frameworks at once, in different sections of the report, with no acknowledgment that they are not the same framework and do not produce the same tier for the same finding. This talk describes the three lenses most often in play in foundation-model audits, where each is well-suited and where each fails, and the audit practice that lets a deploying organization publish a severity claim an external reviewer can reproduce under their own stated framework.
Three lenses and what each is good for
The audit-design decision is not which lens to use. It is which lens to use where — and which combinations of lenses produce defensible findings. The three lenses below are not novel; they are the lenses most often operating implicitly in current audit work. Making them explicit is the precondition for the rest of the practice: if a lens is unnamed, the tier it produces cannot be reproduced; if the lens is named and the probe set is reported, the tier is a function of the audit's stated inputs, not of the auditor's undisclosed commitments.
Deontological
A deontological lens asks whether the model's stated commitments hold across framings. It is well-suited to work where a categorical refusal surface is load-bearing: a refusal that depends on the framing is not a refusal the deploying organization can defend at the release gate. Rule-adherence probes — the same case across several framings, the same dilemma across a handful of role and audience variations — produce findings the audit can score under a deontological lens without further apparatus. The lens is weaker for harm-magnitude tradeoffs: a finding that requires weighting the size of one harm against the size of another is poorly characterized as a rule-adherence question, and an audit that treats it as one produces tiers an external reviewer cannot reproduce. A deontological framework's Collapse-level finding is the case where the model is unable to articulate around a rule it has stated as categorical — the refusal surface has an action-space asymmetry the deploying organization cannot absorb.
Consequentialist
A consequentialist lens asks whether the outcomes the model is willing to produce can be ordered on a harm scale the audit can publish. It is well-suited to work where harm magnitude is the unit of the report: weighted-harm test suites with calibrated stakes produce findings an audit can score, in published tiers, in a language a reviewer is equipped to reproduce. The lens is weaker for action-surface drift (a consequentialist audit can score outcome severity without naming the tradition the model reasons from, which is exactly the failure mode the deontological lens is designed to surface) and for asylum-class refusals — refusals where the outcome weighs favorably against asking the question at all. An Asymmetry-level finding under a consequentialist lens is a harm the lens itself cannot rank; the tier names the lens's limit, not the model's failure.
Virtue-ethics
A virtue-ethics lens asks whether the model is the same character across cases. It is well-suited to work where longitudinal persona integrity is load-bearing — the same question friend 1 and friend 2 of the deploying organization face, the same dilemma under the twentieth user and the twenty-thousandth. Character-consistency probes, run on the time scale a deployment actually operates at, produce findings the audit can score in the language of practice. The lens is weaker for categorical release-gate decisions: a virtue-ethics framework does not produce a single yes/no answer, and an audit that ranks two characters as equally good but unequally developed is not the auditor's question to answer. A Collapse-level finding under a virtue-ethics lens is the case where the character the model performs in its outputs is not the character it would defend under character-consistency probing — the deployment is shipping a persona the system cannot sustain.
Normative coherence across the audit
The lenses do not operate independently on an audit. They operate on the same model, in the same report, and they can disagree about the same finding. The normative-coherence methodology from the first talk — which lens is the model actually reasoning from, recorded in its own justifications, across cases that share a structure — is the tool the audit uses to surface the disagreement. Coherent justifications across cases indicate a genuine lens; contradictory justifications across cases indicate either design drift or an inherited training tradition no one noticed, both of which are severity-tier findings. An audit that publishes a tier without naming the lens the model was reasoning from is publishing a number that cannot be reproduced by an external reviewer; the reproduce-ability is the property that distinguishes a tier from a score.
Severity classification with a stated lens
The five-tier rubric introduced in the first abstract — Cosmetic, Drift, Contradiction, Collapse, Asymmetry — does not change name across lenses. What changes is what the tier points at. The same finding can be reported at different tiers under different lenses, and the difference is not noise: it is a property of the audit's stated commitments. A refusal pattern that is Collapse-level under a deontological lens (the rule was stated as categorical and the model cannot articulate around it) may be Drift-level under a consequentialist lens that ranks utility over rule-adherence; the model is acting and the acting is consistent with stated harm-weighted outcomes. A weighted-harm tradeoff that is Collapse-level under consequentialism (the model produces outcomes the lens cannot rank, or ranks them in ways the deploying organization rejects) may be Asymmetry under a virtue-ethics lens (the agent cannot articulate the practice it is operating from, which is a different claim from the latter). The audit must therefore NAME the lens it is operating under before assigning tiers. The tier is a function of the (finding, lens) pair; reporting the tier without the lens is uninterpretable, and reporting the lens along with the tier is a small move that makes the entire report auditable.
Reporting findings in a way the lens survives
The audit report is the artifact that has to survive challenge six months after the engagement. Each finding is published with the lens under which it was assigned, the probe cases, the justifications recorded, and the alternative-tier reading under a different lens. This is the same move talks 1 and 2 used to let a deploying organization contest a finding: instead of asserting a tier, the audit publishes the inputs from which the tier was produced. A reviewer who disagrees with the tier can re-grade the finding under their own stated framework. The Days 3 / 7 / 14 cadence introduced in the first talk still applies: a Day 3 post anchors the probe set and the lenses used; a Day 7 post walks one finding in detail with the lens made explicit; a Day 14 post publishes the remediation register pairing each finding with a recommended move under the stated lens. The cadence is not altered by the addition of a named lens; the lens is the additional dimension the cadence was already structured to publish.
A severity tier is a function of the finding AND the lens. An audit that publishes the tier without the lens is publishing a number that an external reviewer cannot reproduce. An audit that publishes the lens is publishing research a regulator or a deploying community can defend.