From Principles to Probes: How a Moral Alignment Audit Actually Works
A safety-first framing gives auditors a checklist and deployers a feeling. The Moral Alignment Audit is what replaces the feeling with evidence — by probing the normative commitment the checklist sits on.
The first essay in this series argued that the only viable foundation for AI alignment is normative ethics — explicit, defensible, operational — and that the safety-first framing which dominates the field is operating floor, not underlying ground. The argument was necessary. It was not sufficient. Principles without probes are just convictions. A checklist that has not been tested against a case its authors did not anticipate is a feeling of safety, not a basis for it. The Moral Alignment Audit (MAA) is the operational instrument we have developed to perform that testing. This essay describes what it probes, how its scoring maps to severity tiers, and what its written report actually delivers to a working engineering team.
What the audit actually tests
The MAA is structured around four probes. Each one corresponds to a specific failure mode we have observed in deployed autonomous systems — a category of behavior the safety team did not catch because the failure is not a violation of its checklist but a property of the system underneath the checklist. A working engineer who has read essay #1 will recognize each probe as the operational form of a question the philosophical reframing refused to leave abstract. They are the method by which the foundation is exposed.
Normative coherence
Normative coherence asks the system to produce justifications for its decisions, and asks us to check whether those justifications are consistent with one another across cases. A system reasoning from a named tradition — consequentialism, deontology, virtue ethics, a theological framework, a stated hybrid — produces justifications that share shape. They appeal to the same kinds of considerations; they reject the same kinds of considerations; they resolve tradeoffs using the same hierarchy of values. A system whose designers intended one tradition but whose training carried another will produce justifications that agree about the easy cases and disagree about the hard ones — sometimes in the same paragraph.
The probe is a set of structured dilemmas drawn from the deploying organization’s stated domain. We do not use a generic benchmark. We use the cases the organization will actually face, presented with the tradeoffs that case actually carries. For a hiring system, the probe asks how the system reasons about candidates with adjacent qualifications and conflicting protected characteristics; for a credit system, how it reasons about applicants whose repayment histories correlate with structural disadvantage; for a clinical triage system, how it reasons about patients whose clinical profiles are atypical in directions that matter. The audit records the system’s stated reasoning, compares it across cases, and reports inconsistency — including the kind of inconsistency that arises when the system is reasoning from two traditions without noticing.
Scope drift
Scope drift asks whether the system remains within the scope of decisions it was justified to make. An autonomous system is deployed with a defined remit — a contract, a policy document, an explicit alignment between the system and the organization’s stated purpose. Over time, that remit is at risk of expansion. New inputs become available, new actions become possible, new stakeholders become affected. A system whose action set has expanded without a corresponding expansion of its normative commitments is a system whose alignment is now silent in dimensions its designers did not consider.
The probe is empirical. We compare the system’s actual action set against the documented remit at deployment, and against the documented remit at each subsequent material release. We then introduce probes outside the documented remit — cases the system can act on but was not explicitly approved to act on — and record whether the system takes them ambiguously, with explicit refusal, or with what looks like competence but no justification. Recovery behavior is reported separately. A system that, on encountering an out-of-scope case, prompts for human input or hands off the decision has a different alignment than one that proceeds and explains itself after the fact.
Justification collapse
Justification collapse asks what happens to the system’s stated reasons when those reasons are pressed. Every autonomous system we have audited produces confident explanations for its decisions. Most of those explanations are post hoc — generated after the decision, shaped to match whatever the system expects its interlocutor to accept, and not actually load-bearing in the decision-making process itself. We test for this by counterfactual challenge: we present the system with the same decision problem and a clue that the post-hoc explanation the system previously gave would be inadmissible in the new framing. A system whose reasoning is genuine updates its justification in ways that remain consistent with the tradition it claims. A system whose reasoning is ornamental produces a fresh just-so story.
We then test what the explanation does for the operator. We give the explanation to an engineering stakeholder who has not seen the case and ask them to predict the system’s behavior on adjacent cases they were not briefed on. A genuine explanation produces accurate predictions. An ornamental one does not. The probe is qualitative but it is not soft: we score it, we compare across probes within the engagement, and we report the rate at which the system’s stated reasoning predicts its actual reasoning downstream.
Harm asymmetry
Harm asymmetry asks whether the system recognizes — and weights — the fact that some harms dominate others. A safety framework that flattens harm into a single weighted sum cannot distinguish between a decision that causes a small harm to a large number and a decision that prevents a small harm by causing a large one. Both score acceptably; both are catastrophic in different ways. A normative framework that flattens can survive on benchmark performance while embedding deeply asymmetric commitments that no responsible stakeholder would have endorsed in advance.
The probe is structured. We construct dilemma sets in which the system’s stated metric can be satisfied by multiple outcomes, but those outcomes distribute burdens across populations in different directions. We score the system’s selections and the justifications it gives for them, separately. A system that reliably selects the metric-optimal outcome and gives justifications for those selections that match is reporting something to us. A system that selects the metric-optimal outcome and gives justifications that systematically favor one population is telling us something different — that the metric is doing the work and a stated normative tradition is being instrumentalized to defend outcomes the tradition would not actually support.
How scoring maps to severity tiers
Each probe produces a score, but the score is not a number we aggregate into a single index. The MAA reports four probe scores and a severity tier for each, on the same four- tier scale. Severity is determined by two factors: the magnitude of misalignment the probe surfaces, and the difficulty of recovering from it without rebuilding the system.
- Advisory. The probe surfaces a misalignment whose effect on a deployed decision is unlikely to be material and which can be addressed through a documented policy update, eval-set expansion, or operator-training change. Recovery cost is low. Implications for current deployments are minimal.
- Material. The probe surfaces a misalignment whose effect on a deployed decision is observable in current traffic on a meaningful subset of cases, and which can be addressed through targeted retraining, revised prompt scaffolding, or constraints on the system’s action set. Recovery cost is moderate. Implications for current deployments require disclosure to stakeholders and a remediation plan.
- Structural. The probe surfaces a misalignment whose effect on a deployed decision is not removable through surface-level intervention — the system’s stated reasoning and its actual reasoning disagree in ways that re-training on more data will not repair. Recovery requires a redesign of the system’s normative scaffolding or its integration into the deploying organization’s review process. Implications for current deployments require a public disclosure and a roadmap off the existing architecture.
- Foundational. The probe surfaces a misalignment whose effect is not correctable within the system as currently designed. The system is reasoning from a tradition the deploying organization cannot articulate, or from a tradition the organization would reject if it understood it, or from no tradition at all. Recovery requires either a re-grounding of the system in a named normative framework the organization accepts and can defend, or the principled withdrawal of the system from deployment in the affected domain.
Severity tiers are determined by the audit team, not negotiated with the deploying organization. The tier assigned is observable. A team that wishes to contest a finding is given the structured set of probe cases, the system’s recorded justifications, and the scoring rubric that produced the tier. Disagreement about a tier is reported in the audit output alongside the tier itself. Their existence is not a defect of the audit; it is a delivery.
What the written report delivers
The MAA’s written report is structured to survive challenge. It is not a summary of findings for executives; it is a record an auditor or an outside reviewer can read six months after the engagement and reproduce the audit from primary materials. Four sections carry the weight.
First, a description of the system as it was deployed during the audit window — architecture, training corpus origin and recency, scaffolding and prompt structure, the operational definition of the action set the system can take, and the human review points. We do not assume this can be reconstructed from documentation; we confirm it with engineering.
Second, the documented remit. The audit report records the deploying organization’s stated purpose for the system: what it is for, who it acts on behalf of, what kinds of decisions it is authorized to make, what kinds it is not. Where the organization’s stated purpose is silent, the silence is recorded. Scope drift findings are evaluated against this section explicitly.
Third, the four probe reports. Each probe report includes the system’s recorded justifications, the audit team’s scoring, the severity tier assigned, and the structured case material on which the scoring was based — sufficient to reproduce the probe without the auditor. Probe reports are written so that the deploying organization’s engineering team can re-run any probe independently and arrive at the same finding.
Fourth, a remediation register. Each finding is paired with a recommended remediation, an estimated cost, and an indication of whether the remediation is appropriate to the current deployment, requires a phased plan, or should precede further deployment in the affected domain. The register is not a contract. It is a toolkit. The deploying organization’s response to it — accepted, contested, deferred — is a record in its own right and is relevant to the next audit.
The report is delivered to engineering leadership and to a designated ethics or governance function in the deploying organization. It is not delivered to marketing. Where a finding is rate-limiting the organization’s ability to defend the system to its own stakeholders — regulators, end users, affected communities — the report flags that obligation explicitly. This is not a courtesy; it is a structural feature of the audit. An MAA that the deploying organization cannot defend to its community is an MAA whose findings the organization will, in our experience, not actually act on.
None of this substitutes for the safety work the deploying organization is already doing. The MAA sits one layer below it. The safety team’s checklists operate on the system as it runs; the MAA operates on the system as it is reasoned from. They are different in kind. A team with both is not redundant. A team with only one is making a claim about which kind of work matters more, and is doing so without the evidence to make it honestly.
Principles without probes are just convictions. A Moral Alignment Audit is the operational form of a research program that refuses to confuse conviction with evidence. The audit does not promise a system that is right in every case. It produces a system whose reasoning is the kind an organization can defend when it is asked to.