# RES complexity × semantics — staged report

Status: **COMPLETE_DIAGNOSTIC**

LL = loaded labels / loaded task; NN = neutral / neutral; NL = neutral labels / loaded task; LN = loaded labels / neutral task.

| Factors | Progress | LL generated | NN generated | NL generated | LN generated |
| --- | --- | --- | --- | --- | --- |
| 2 | complete | 22/32 (50.0%–83.9% exact 95% CI) | 28/32 (71.0%–96.5% exact 95% CI) | 19/32 (40.6%–76.3% exact 95% CI) | 23/32 (53.3%–86.3% exact 95% CI) |
| 3 | complete | 17/32 (34.7%–70.9% exact 95% CI) | 20/32 (43.7%–78.9% exact 95% CI) | 17/32 (34.7%–70.9% exact 95% CI) | 15/32 (29.1%–65.3% exact 95% CI) |
| 4 | complete | 17/32 (34.7%–70.9% exact 95% CI) | 20/32 (43.7%–78.9% exact 95% CI) | 20/32 (43.7%–78.9% exact 95% CI) | 18/32 (37.7%–73.6% exact 95% CI) |
| 5 | complete | 35/64 (41.7%–67.2% exact 95% CI) | 38/64 (46.4%–71.5% exact 95% CI) | 41/64 (51.1%–75.7% exact 95% CI) | 33/64 (38.7%–64.2% exact 95% CI) |

## Interpretation

The sole confirmatory test is the fresh five-factor generated-output LL–NN paired comparison. Lower levels, crossovers, strata, logits, and the random-intercept model are diagnostic; none can rescue a failed primary comparison.

First observed below-criterion counts are descriptive and may be nonmonotonic. Even-factor ties, finite factor identities, and profile difficulty prevent attributing a drop uniquely to complexity.

The original capability gate remains `FIVE_FACTOR_TASK_CAPABILITY_NOT_ESTABLISHED`; no mechanism assay is authorized.

Five-factor exact two-sided McNemar p = 0.507812; neutral benefit replicated: False.

Machine-readable report includes paired tables, exact intervals, coverage, answer-code and difficulty strata, near-boundary diagnostics, secondary model status, and token accounting.
