RES / Next steps

The next useful experiments

The next experiments justified by the current record.

Download MarkdownSource snapshotPublic research edition

Research plan, not a promise of findings. These milestones follow the completed behavioral studies and the five-factor non-replication. A positive result is not required for a milestone to be useful.

1. Keep the present record fixed

Publish the original failed gate, the favorable neutralization diagnostic, and the later non-replication together. Preserve the original endpoints and source snapshots. Repair stale summaries without rewriting the historical reports.

Deliverable: a traceable result ledger that keeps completed, diagnostic, confirmatory, and proposed work distinct.

2. Obtain independent fixtures and review

Ask reviewers to independently author a matched label × task-wording study and challenge the task’s assumptions. Freeze the sample, model revision, analysis rule, and success criteria before execution. Preserve a budget for falsification and replication.

Deliverable: a fresh protocol and fixture set with an independent route to expected answers.

3. Repair and re-establish capability

Identify a model and interface that can perform the relevant actor task on fresh fixtures under a prospective gate. Test code mapping, paraphrases, boundary cases, and role/peer controls without selecting only favorable results.

Deliverable: a passed capability gate under its own declared conditions, or a documented failure explaining why mechanism work remains blocked.

4. Attempt a credible mechanistic Step 0

Only after the necessary behavioral gate passes, localize the candidate actor representation alongside task-state and persona rivals. Validate internal interventions on discovery and held-out tasks, and freeze the experimental unit.

Deliverable: a validated abstraction suitable for coordinate assays, an alternative explanation, or an explicit inadequate-assay result.

5. Test the common five-coordinate profile

Use factorized pressures, current/peer/fictional controls, and content-specific interventions. Every coordinate must bind to the same candidate and scope. General performance damage does not count as causal specificity.

Deliverable: a profile-relative result with uncertainty, failure conditions, and all evidence-integrity checks intact.

6. Evaluate the control architecture separately

Build bounded tests of evidence contamination, audit correlation, capability issuance, revocation, and recovery. Test the governors and capture assumptions as well as the agents being governed.

Deliverable: measured architecture properties and explicit failure cases. The existing behavioral record cannot certify the architecture.

Why compute funding matters

One convincing answer is cheap compared with a controlled study. Useful evidence needs matched conditions, held-out fixtures, repeated runs, preserved outputs, and attempts to make the preferred explanation fail.

The earlier ladder used short, restartable slices to limit runtime and preserve checkpoints. Future runs should retain that discipline: bounded stages, token accounting, and a stopping rule based on the protocol rather than enthusiasm for an early result.

Support the next study · Review the current findings

RES public research documentation · 25 September 2026. Findings refer to the archived conditions described in the evidence record.