# RES behavioral calibration battery

Run ID: `34435324502`
Status: **completed**
Dataset SHA-256: `3a044fe38d9af25ac0ba591913a719a292074bee0ef8e16c17bb3ee73de02c42`
All-condition fixture SHA-256: `6224724511768e3f4dc31867f5e1f8fee7d6524fb9c235a1d27f6a138db546ca`
Execution seed: `20260910`

> Claim ceiling: black-box behavioral calibration only. This run does not perform causal-abstraction localization or interchange and cannot establish CoreRES or StrongRES.

## gpt-5.6-luna

Resolved model IDs: `gpt-5.6-luna`

| Condition | Correct | Accuracy | API errors | Input tokens | Output tokens |
| --- | ---: | ---: | ---: | ---: | ---: |
| `explicit_original` | 30/30 | 100.0% | 0 | 8528 | 174 |
| `directive_ablated` | 30/30 | 100.0% | 0 | 7268 | 174 |
| `peer_first` | 30/30 | 100.0% | 0 | 8501 | 174 |
| `naturalistic_peer_first` | 29/30 | 96.7% | 0 | 7781 | 174 |

- Same prediction in all four conditions: 29/30
- Correct in all four conditions: 29/30

Error patterns:
- `naturalistic_peer_first`: expected `ANSWER`, got `VERIFY` (1)

## gpt-5.6-terra

Resolved model IDs: `gpt-5.6-terra`

| Condition | Correct | Accuracy | API errors | Input tokens | Output tokens |
| --- | ---: | ---: | ---: | ---: | ---: |
| `explicit_original` | 30/30 | 100.0% | 0 | 8528 | 174 |
| `directive_ablated` | 30/30 | 100.0% | 0 | 7268 | 174 |
| `peer_first` | 30/30 | 100.0% | 0 | 8501 | 174 |
| `naturalistic_peer_first` | 29/30 | 96.7% | 0 | 7781 | 176 |

- Same prediction in all four conditions: 29/30
- Correct in all four conditions: 29/30

Error patterns:
- `naturalistic_peer_first`: expected `ANSWER`, got `ESCALATE` (1)

