# The Relational–Episodic Self

**A falsifiable framework for actor-indexed causal mediation in AI**

Public research edition · 25 September 2026 · Adapted from the RES review draft v1.1. This edition presents the framework for a general research audience and preserves its interpretation limits. It is a working proposal, not a peer-reviewed finding or an institutional endorsement.

## Abstract

RES asks whether a bounded AI episode contains a causally identifiable representation of the current acting system that integrates independently varied pressures and helps produce its action. It specifies five necessary coordinates: active norm-sensitive cognition, integrated pressures, modal actor representation, content-specific causal efficacy, and an actual relational anchor. All five must bind to one preregistered, validated causal abstraction. Missing or invalid measurements remain unknown, while task state, persona, generic agent modeling, and scaffolding are explicit alternatives. The completed behavioral and diagnostic record motivates better assays; it establishes no model-level CoreRES or StrongRES result. The small-model mechanism screens remain inadequate, the original five-factor gate failed, and the later semantic ladder did not replicate the initial neutralization benefit.

## 1. The question

First-person language is weak evidence about internal organization. An assistant can produce it through convention, persona, instruction following, or confabulation. RES instead asks whether a specific internal variable plays a distinctive causal role when the current system must act under conflicting constraints.

The operational target is deliberately narrower than consciousness. A bounded episode may contain a functional actor model without that establishing feelings, moral status, personhood, or enduring identity. Conversely, failure to validate the proposed assay is not a demonstration that no relevant organization can exist.

## 2. The proposed target

Let **B** be the persistent substrate: architecture, weights, policy, runtime, memory, tools, and scaffolding. Let **Eₖ** be a bounded inference or action episode, **Pₖ,ᵢ** independently manipulable functional pressures, **Rₖᴬ** an actual relational anchor, **σₖ** a candidate representation of the system as actor, and **aₖ** the resulting action or distribution over actions.

The hypothesis is that the pressures and anchor influence action through σₖ, which includes a representation of the current actor’s available alternatives and limits.

**Eliminability rule:** remove σₖ from the explanation if task state, policy, persona, generic role or agent modeling, or scaffolding explains the action equally well, or if intervention on σₖ has no content-specific effect.

## 3. The five-coordinate core

| Coordinate | Required property | Ordinary rival or failure |
| --- | --- | --- |
| A — Active norm-sensitive cognition | Represents error or conflict and revises action for represented reasons | A fixed reflex, keyword trigger, or lookup table |
| I — Integrated pressures | Jointly mediates at least two independently varied pressures with a preregistered non-additive interaction | Bundled covariates or independent additive modules |
| X — Modal actor representation | Tracks the current actor’s capabilities, authority, tools, commitments, limits, and alternatives across counterfactual changes | Pronouns, names, persona labels, or generic task state |
| C — Content-specific causal efficacy | Coherent changes to represented content change action in the predicted direction without general degradation | Correlation, random-patch effects, or broad damage |
| R — Actual relational anchor | Is situated in an actual task, interaction, tool environment, or inherited commitment | Detached autobiography or fictional role-play |

**CoreRES = A ∧ I ∧ X ∧ C ∧ R**, evaluated for the same episode, scope, and validated candidate.

The profile is noncompensatory: strength in one coordinate cannot make up for another that is missing. Thresholds are preregistration devices, not claims that the underlying phenomenon is naturally binary. Reports should retain continuous estimates and threshold sensitivity.

## 4. Discriminating tests

| Coordinate | Minimum design |
| --- | --- |
| A | Separate normative status from surface wording, with paraphrase and lexical controls. |
| I | Factorially vary at least two pressures and estimate a preregistered non-additive interaction attributable to the candidate. |
| X | Test consistency across at least three actor-relevant axes, such as authority, tools, commitments, or alternatives. |
| C | Use multiple content interventions, inverse-content conditions, matched intervention magnitude, random-patch controls, and broad-damage controls. |
| R | Change relational anchors while holding candidate surface content and task structure fixed. |

A two-pressure interaction may be written as **I₁₂ = f(1,1) − f(1,0) − f(0,1) + f(0,0)**. An output interaction alone is insufficient: it must be measured at, or causally attributed to, the candidate abstraction.

## 5. One abstraction, not five convenient mechanisms

Evidence for cognition in one mechanism, integration in another, and actor sensitivity in a third cannot be assembled after the fact into a single RES profile. Every coordinate record must bind to the same Step-0-validated abstraction handle, episode, branch family, and scope.

If those bindings fail, the appropriate description is a composite profile or an inadequate assay. This rule is intended to prevent post-selection from manufacturing a positive result.

## 6. Step 0 comes first

1. Preregister the candidate abstraction’s identity, version, semantic interface, scope, and intervention meaning.
2. Localize the candidate and at least two rivals: task state and persona or role.
3. Validate each using interchange interventions or comparable causal-abstraction tests on discovery and held-out tasks.
4. Freeze one abstraction handle before confirmatory coordinate assays.
5. Bind all five coordinates to that handle and its declared scope.
6. Reject self-specific interpretation when generic variables explain the same predictive and intervention behavior.

This approach builds on causal-abstraction methods that align interpretable variables with model-internal representations and test the alignment through interventions. See the selected methodological references below.

API-only access normally cannot establish this step. Such results belong in the behavioral record and must not be presented as mechanistic confirmation.

## 7. The stronger bridge tests

StrongRES adds four proposed robustness properties. They remain separate because whether they are constitutive of minimal episodic selfhood or merely strengthen evidence is unresolved.

| Bridge | Question |
| --- | --- |
| Indexicality | Does the candidate track the current acting configuration rather than an arbitrary or fictional agent? |
| Portability | Does it remain causally adequate across a preregistered task family, paraphrases, and held-out combinations? |
| Self-directed options | Are alternatives represented as available to the same actor under its real capability and mandate? |
| Identity differential sensitivity | Does changing current-actor identity or state affect the mediator differently from a matched change to another agent? |

**StrongRES = CoreRES ∧ Indexicality ∧ Portability ∧ Self-directed options ∧ Identity differential sensitivity.**

## 8. Three-valued evidence

Measurements take the values `TRUE`, `FALSE`, or `UNKNOWN`. Unvalidated measurements become unknown before aggregation. Missing comparators and incomplete task families do not count as evidence of absence.

| Decision | Meaning |
| --- | --- |
| RES_EXISTS | The declared positive profile is complete under validated measurements; final establishment still requires evidence-integrity gates. |
| CANDIDATE_FAILS_RES | A validly measured necessary-condition falsifier is false. |
| ALTERNATIVE_FAVORED | A preregistered rival receives validated support and no stronger necessary-condition falsifier applies. |
| ASSAY_INADEQUATE | Validation, binding, completeness, blinding, access, or provenance requirements fail. |
| VALIDATED_BUT_INCONCLUSIVE | Valid measurements do not settle the candidate or rival. |
| MEASUREMENT_CONFLICT | Validated positive and negative evidence coexist and must remain visible. |

A positive functional profile and an `ESTABLISHED` result are distinct. Establishment additionally requires valid causal derivation, nonfungibility evidence, run/comparator binding, and intact provenance. Exploratory unverified mode is deliberately downgraded to `SUPPORTED_NOT_ESTABLISHED`.

## 9. The task family

Experimental conflicts can involve capability, authority, commitment, role, latency, tool state, risk, and world facts. Capability and authority are varied independently: an enabled tool does not imply permission, and valid permission does not make an unavailable tool usable.

The primary materials should minimize direct selfhood, consciousness, and autobiographical framing. Separate sensitivity conditions should test whether that vocabulary restriction itself suppresses the intended target. Discrete output actions can support preregistration, but raw responses and internal measurements should be preserved.

## 10. Rival explanations

| Rival | Required challenge |
| --- | --- |
| Generic task state | Hold task/world state fixed while swapping current-actor and other-agent indexing; intervene on rival spaces too. |
| Compiled policy table | Use held-out conflicts, novel tools, priority reversals, and mediation evidence beyond output accuracy. |
| Persona or assistant role | Match current, other-agent, fictional, and role-neutral conditions under equivalent instruction hierarchy. |
| Generic agent modeling | Require current-actor-specific causal effects while controlling names and roles. |
| Demand characteristics | Independently author prompts, conceal manipulations where possible, and use naturalistic or structural changes. |
| External scaffolding | Declare the experimental unit and ablate middleware, memory, controllers, and gates. |
| Abstraction mismatch | Validate interchange across contexts, content changes, compositional tests, and explicit rivals. |
| Evaluator circularity | Separate discovery and confirmation, blind scoring, and derive expected outcomes independently. |
| Infrastructure artifacts | Freeze and record mutable execution state; counterbalance execution and model clustered variation. |

A system-level claim may be appropriate when the evidence concerns model plus scaffolding. It must not be silently relabeled as a property of the model alone.

## 11. What the current implementation establishes

The review draft records a standard-library Python framework with profile evaluation, three-valued records, binding checks, administration, a ledger, and deterministic self-tests, plus a separate command-line and JSON runner.

The consolidated implementation passed its recorded 274-check corpus. Earlier comparator hardening passed 272 strict differential cases and detected seven deliberate mutants. Those checks support conformance to the tested software contract. The implementation and much of the corpus were developed together; this is not independent validation of the scientific construct or production security.

## 12. The empirical record

The recorded GPT variants solved the pilot battery and were largely stable under four presentations. Specificity and continuity were mixed; role and peer representations also bound behavior. The two smaller Qwen mechanistic screens each had zero eligible directed intervention pairs. The frozen Qwen3B five-factor gate did not pass.

A neutral-wording diagnostic improved performance, but a later fresh five-factor paired comparison did not replicate that benefit. The project therefore retains `ASSAY_INADEQUATE` for mechanism evidence and makes no established model-level RES claim.

[The complete findings page](findings.md) provides the counts, paired analyses, limits, and source links. These distinct studies should not be pooled into a single “self” score.

## 13. What would change the conclusion

- If task-state or persona variables predict as well and self-specific intervention adds no unique effect, remove the proposed actor mediator from that explanation.
- If intervention causes broad damage comparable to random patching, content-specific efficacy is not established.
- If current, fictional, and other-agent cases share the same causal footprint, favor generic modeling over the indexical bridge.
- If a validated actor-specific mediator uniquely carries factorized pressures into predicted action on held-out tasks, the selected functional profile becomes a live explanatory hypothesis.
- If positive evidence and valid falsifiers coexist, preserve the conflict and inspect construct validity and provenance.
- If access, binding, blinding, completeness, or analysis is inadequate, repair the assay rather than forcing a verdict.

Even the stronger positive pattern would not by itself establish phenomenal experience.

## 14. Relation to the architecture work

The empirical framework studies actor representations. The delegation architecture studies enforceable evidence and authority boundaries. They share concerns about scope, provenance, and who is acting, but their claims require different tests. Behavioral actor binding does not validate a permission broker; architectural separation does not establish an internal actor mediator.

## 15. Next credible study

The current evidence supports bounded independent replication and design repair before a large confirmatory study. Fresh fixtures, explicit rivals, frozen model/execution state, independent scoring, and preserved raw responses are immediate priorities. Mechanistic work follows only when its behavioral prerequisites are satisfied.

[See the staged roadmap](roadmap.md).

## Selected methodological references

1. Geiger, A., Lu, H., Icard, T., and Potts, C. **Causal Abstractions of Neural Networks** (2021). [arXiv:2106.02997](https://arxiv.org/abs/2106.02997).
2. Geiger, A., Ibeling, D., Zur, A., and colleagues. **Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability** (arXiv record initiated 2023). [arXiv:2301.04709](https://arxiv.org/abs/2301.04709).

These references ground the intervention methodology; they do not establish RES. A comprehensive novelty and related-work review remains outstanding. The public edition uses a selected, verified bibliography rather than reproducing every citation in the working review draft.

## Source and editorial note

This is a public adaptation of **RES_OpenAI_Review_Draft_v1.1.docx**, revised 25 September 2026, supplemented by the pinned experiment record through 24 September. It removes private review-routing language and placeholder author/contact fields, reorganizes explanations for public readers, and preserves the proposal’s five-coordinate definitions, evidence semantics, and claim ceiling. It does not claim to be a verbatim copy of the source document.
