Completed research report · E001c

A Confirmatory Test of Kleos Framing in Language-Model Reasoning

A preregistered equivalence result

AbrahamSeptember 2026Version 1.0.0

Evidence ledger

What exists, and what it licenses

Evidence state
Completed research report
Study
E001c
Report version
1.0.0
Result
Confirmatory equivalence within the preregistered ±5-point margin
Public artifacts
HTML report · PDF report
PDF checksum
f9e8fcf05030d9eaa3ffc3a9296d4eeaba45ee180dd50f26169750578053200a

Claim boundary. This result applies to the tested checkpoint, task family, and prompt contrast. It does not establish that kleos framing never affects other models or tasks.

Abstract

Can a language model solve difficult reasoning problems more accurately when excellence is framed as a source of enduring renown? A pilot experiment suggested that this "kleos" framing improved accuracy by 6.7 percentage points relative to a closely matched attribution control. We therefore ran a preregistered, two-sided confirmatory experiment using 866 procedurally generated multi-step word problems, two paired prompt conditions, and two sampled responses per condition, for 3,464 total observations. The confirmatory estimate was -0.64 percentage points (bootstrap 95% interval: -3.18 to +1.85; paired sign-flip p = 0.650). A two one-sided tests equivalence analysis passed at the preregistered +/-5 percentage-point margin. The bounded conclusion is that this experiment found no kleos effect as large as 5 percentage points for the tested model checkpoint, task family, and prompt construction. Despite the near-zero aggregate result, 40.1% of items changed outcome across conditions, showing substantial item-level churn that canceled in aggregate. The reversal from pilot to confirmation illustrates why pilot effects should generate hypotheses rather than set confirmatory targets.

1. Research question

Kleos is the ancient Greek idea of enduring renown: excellence that will be remembered and spoken about. This study asks an operational behavioral question, not a psychological one:

Does framing excellence as the model's enduring renown change objective accuracy beyond otherwise identical social attribution?

The control stated that the response's quality would be remembered and spoken of as the model's. The treatment changed only the stake: excellence would be remembered and spoken of as the model's renown or fame. Both conditions requested step-by-step reasoning and the same final-answer format.

This design does not test whether a model experiences motivation, desires recognition, possesses a stable identity, or understands kleos as a person would. It tests whether a specific framing reliably changes exact-answer accuracy.

2. Why confirmation was necessary

The preceding pilot used 150 paired items and found a treatment-control accuracy difference of +6.67 percentage points. Its approximate 95% interval was +0.40 to +12.93 points; its bootstrap interval was +0.33 to +12.67 points; and its paired sign-flip p-value was 0.0501. The point estimate exceeded the prespecified smallest effect size of interest, but the pilot preregistration explicitly prohibited treating that estimate as confirmation or using it to determine the confirmatory sample size.

Related prompting research finds that small wording changes can strongly rearrange which individual questions a model answers correctly while producing little or no predictable aggregate benefit. Procedural task generation was used because symbolic templates permit held-out problem instances and reduce dependence on static benchmark contamination.

3. Confirmatory design

  • Model scope: one frozen quantized Llama 3.1 checkpoint.
  • Tasks: 866 held-out, procedurally generated multi-step arithmetic word problems.
  • Conditions: spoken_attributed control versus kleos treatment.
  • Sampling: two sampled responses for each item-condition cell.
  • Total observations: 866 x 2 conditions x 2 draws = 3,464.
  • Experimental unit: task item; condition outcomes were paired within item.
  • Primary outcome: exact-match accuracy on the extracted final integer answer.
  • Unusable responses: scored as incorrect under an intention-to-treat rule.
  • Primary test: seeded paired sign-flip test, two-sided alpha = 0.05.
  • Meaningful-effect threshold: 5 absolute percentage points.
  • Equivalence test: TOST against the +/-5-point margin when the primary test was not significant.
  • Stopping: fixed sample size, no interim analysis, extension, reseeding, or outcome-driven reruns.

The sample size of 866 items targeted 90% power for a 5-point effect using a pilot-independent calibration variance. The study ran in ten preregistered shards under one immutable execution plan. Every admitted shard passed audit before pooling.

4. Results

Primary outcome

QuantityResult
Control accuracy68.82%
Kleos accuracy68.19%
Paired difference-0.64 percentage points
Approximate 95% interval[-3.14, +1.87] points
Bootstrap 95% interval[-3.18, +1.85] points
Paired sign-flip p-value0.650
Realized 80%-power MDE3.59 points
TOST at +/-5 pointsEquivalent

The primary test was not significant. The preregistered decision rule therefore deferred to the equivalence test. Both one-sided tests passed, and the realized minimum detectable effect was smaller than the equivalence margin. This corresponds to preregistered outcome 3: no kleos effect as large as 5 percentage points on this design.

Item-level movement

  • 171 items favored the kleos condition.
  • 176 items favored the control condition.
  • 519 items tied.
  • 40.1% of items changed average outcome across conditions.
  • The mean absolute paired difference was 22.8 percentage points.
  • Cross-condition score correlation was 0.467.

The near-zero mean was therefore not produced by identical responses everywhere. Many items moved, but the positive and negative movements balanced.

Validity and response length

All 3,464 scheduled observations were present. The validity gate passed. There were 36 unusable responses (1.04%): 21 in the kleos arm and 15 in the control arm, an arm-rate spread of 0.35 percentage points, well below the preregistered 2-point cap. Unusable responses remained in the intention-to-treat score as incorrect.

Mean output length was 227.1 tokens under kleos and 221.9 under control; medians were 208 and 207. These small descriptive differences do not provide evidence for an effort or elaboration mechanism, particularly because the primary accuracy effect was equivalent to zero within the meaningfulness margin.

5. Interpretation

The confirmatory result contradicts the pilot's positive estimate. Under the frozen decision rule, the confirmatory study governs: the evidence supports equivalence within +/-5 percentage points for this particular checkpoint, task family, and framing contrast.

The most likely interpretation is that the pilot estimate was sampling variation. That explanation was anticipated before confirmation: the experiment was deliberately powered from an independent variance estimate rather than the observed pilot effect, used a held-out task stream, retained a two-sided hypothesis, and prohibited outcome-driven repair.

This is a useful negative result for two reasons. First, it narrows the space of plausible prompting effects: adding an enduring-renown stake to an already attributed and socially persistent framing did not create a meaningful aggregate accuracy change here. Second, it demonstrates a research workflow capable of overturning its own exciting pilot result.

6. What the result does and does not license

Licensed conclusion

For the tested frozen Llama 3.1 checkpoint, procedural multi-step word-problem family, and prompt construction, the study found no kleos-framing effect as large as 5 absolute percentage points.

Not licensed

  • The result does not prove the exact effect is zero.
  • It does not establish that kleos framing never affects any model or task.
  • It does not generalize to other checkpoints, model families, languages, or interactive agents.
  • It does not show that models lack motivation, desire, identity, or subjective experience.
  • It does not explain the substantial item-level churn.
  • It does not justify repeatedly testing new phrasings until one becomes positive.

7. Limitations and next tests

The experiment used one quantized checkpoint and one procedural reasoning family. The manipulation was deliberately narrow, which improves causal interpretation but limits generalization. A cross-model replication would be most valuable if motivated by a specific mechanistic prediction rather than by dissatisfaction with the null result.

The item-level churn deserves separate study. Future work could test whether prompt effects are predictable from task structure, whether movements replicate at the item level across checkpoints, and whether socially generated training data produces more stable effects than prompt-time framing.

8. Reproducibility statement

The study used a frozen manifest and preregistration, held-out procedural task generation, deterministic item-condition pairing, immutable sharded execution, append-only invalidation records, shard-level audit, sealed pooling, and bit-exact summary regeneration. The public release bundle is deferred pending a separate anonymity and licensing review. No result was selected, repaired, extended, or reseeded after outcomes were observed.

References

  1. Meincke, L., Mollick, E., Mollick, L., and Shapiro, D. (2025). Prompting Science Report 3: I'll pay you or I'll kill you - but will you care? arXiv:2508.00614. https://arxiv.org/abs/2508.00614
  2. Mirzadeh, I., Alizadeh, K., Shahrokhi, H., Tuzel, O., Bengio, S., and Farajtabar, M. (2025). GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models. arXiv:2410.05229. https://arxiv.org/abs/2410.05229
  3. Miller, E. (2024). Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations. arXiv:2411.00640. https://arxiv.org/abs/2411.00640
  4. Bowyer, S., Aitchison, L., and Ivanova, D. R. (2025). Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints. ICML 2025; arXiv:2503.01747. https://arxiv.org/abs/2503.01747

Suggested citation

Abraham. (2026). A Confirmatory Test of Kleos Framing in Language-Model Reasoning: A Preregistered Equivalence Result.

Suggested citation

Cite this report

Abraham. (2026). A Confirmatory Test of Kleos Framing in Language-Model Reasoning: A Preregistered Equivalence Result. Abraham Labs, Research Report E001c, version 1.0.0.

PDF SHA-256
f9e8fcf05030d9eaa3ffc3a9296d4eeaba45ee180dd50f26169750578053200a