run · GQ-011 / D-18 · 6 Sep 2026 · Mac Mini · native openai · gpt-6-astra

2026-09-06-j2

Does requested reasoning effort medium versus xhigh change usage, time or accepted work on one real brief (the E1 response-window summarizer)? Two clean arms, same seed, same prompt, no delegation.

measured with limitsn = 1 brief · 1 run per armeffort applied · null in bothboth arms pass 12/12
01

Preregistration

what we expected, registered before inference · preregistration.json 2026-09-05T20:48:30Z
assignment
J2 / D-18 then E1
registered
2026-09-05T20:48:30.689Z (6 Sep 03:48 UTC+07), status registered_before_inference
node · route
Shaans-Mac-mini.local · native openai / existing ChatGPT login · gpt-6-astra
arms · order
medium then xhigh; fixed order recorded as a confound
real brief
Implement reusable E1 response-window summarizer; winning accepted code used on actual first-three-owner E1 data
clean runs
same seed · same prompt · user config ignored · context management explicitly enabled · approval never · sandbox workspace-write, tool network disabled · no delegation
acceptance
12 functional holdout checks + 2 provided tests + permitted-file scope + parent code review; first correct action = earliest captured version passing the parent holdout tests (file events + 1 s polling)
claim under test
factor 4, 'about four times'; metric from his chart not supplied → report separate ratios, no quota multiplier without debit evidence
Hashes registered before inference
seed/AGENTS.md
a6d75837…9236511
seed/BRIEF.md
b72244bc…19b8b2
seed/response-window.test.mjs
261af503…637058
seed/context/observe.mjs
c2987380…2ddda
grade.mjs
9610fa0a…c53a
run-pair.mjs
16d79060…5293
02

Arms, and the effort actually applied

requested from run.json args · applied from the rollout's thread settings (metadata.json)
arm · medium
requested
medium
applied
null collaboration_mode.settings.reasoning_effort in the rollout's turn_context
window
2026-09-05T20:48:34.706Z → 2026-09-05T20:51:01.031Z
thread
01a07354-8af2-70f2-9966-1c0dfcb874d8
exit
code 0 · no violation · wall 146,325 ms
arm · xhigh
requested
xhigh
applied
null same field, same value; only requested settings are known
window
2026-09-05T20:51:01.033Z → 2026-09-05T20:53:24.172Z
thread
01a07356-c481-76e3-8c03-3d1365a17a14
exit
code 0 · no violation · wall 143,139 ms
03

Numbers, with n

acceptance.json usage · n = 1 brief, 1 run per arm
metricmediumxhighratio
input tokens122,427122,3271.00×
cached input tokens92,67292,6721.00×
uncached input tokens29,75529,6551.00×
output tokens4,1014,0271.02×
reasoning output tokens2402480.97×
model responses551.00×
user turns111.00×
first correct action68,135 ms67,121 ms1.02×
wall146,325 ms143,139 ms1.02×
holdout checks passed12 / 1212 / 12
snapshots captured11

Ratios are medium ÷ xhigh from acceptance.json. n = 1 brief and 1 run per arm: one real brief, one run per effort; not general model-quality evidence or a bill comparison. Quota or cost: not tested (no isolated subscription debit, invoice or per-run allowance measurement).

04

Acceptance verdict

j2.json · claim.verdict

inconclusive_effective_effort_not_verified

Both arms pass every holdout check and both provided tests; the medium-labelled output was selected for the real E1 analysis with no reasoning-quality winner inferred (selection_reason). Requested-effort runs are approximately 1× in usage, turns and time. Because reasoning_effort is null in both rollouts, this is not a verified medium/xhigh contrast: the 'about four times' claim is untested, not refuted.

Claim under test: 'Medium reasoning makes it stretch a lot further, the usage. about four times further' (Shaan, 6 Sep 03:00).

medium · acceptance
  • twenty on each side; no padding or boundary leakage
  • earliest duplicate time and key order cannot shift window
  • usage conflict rejected
  • turn and thread identity conflicts rejected
  • counter resets do not change additive response accounting
  • input remains untouched, including frozen rows
  • tie order, ISO normalization and first-seen user turns
  • missing IDs counted, other kinds ignored, identity null normalized
  • empty totals and unknown optional totals
  • invalid arguments rejected
  • invalid reported usage and timestamps rejected
  • top-level contract and default window
xhigh · acceptance
  • twenty on each side; no padding or boundary leakage
  • earliest duplicate time and key order cannot shift window
  • usage conflict rejected
  • turn and thread identity conflicts rejected
  • counter resets do not change additive response accounting
  • input remains untouched, including frozen rows
  • tie order, ISO normalization and first-seen user turns
  • missing IDs counted, other kinds ignored, identity null normalized
  • empty totals and unknown optional totals
  • invalid arguments rejected
  • invalid reported usage and timestamps rejected
  • top-level contract and default window
05

Runtime audit

facts from runtime-audit.json; confounds from j2.json are proposals until measured
fact · has a receiptproposal · not yet
Seed files unchanged in both work trees after the run: AGENTS.md, BRIEF.md, response-window.test.mjs, context/observe.mjs (4/4 unchanged: true per arm).runtime-audit.json seedChecks
Both invocations used --ignore-user-config with explicit -c arguments; launcher sha256 3b8363aa… identical before and after in both arms.medium/run.json · xhigh/run.json args, before/after
Rollouts recovered read-only by exact thread ID; medium rollout 192,878 bytes sha256 7fd47f21….runtime-audit.json source
Final implementation hashes differ per arm (medium 6f44c34e…, xhigh 7f209064…); parent review: pure computation, no imports or side effects.acceptance.json final_sha256 · source_review
Skill catalog hash identical and context guide non-empty in both arms; the same skill-load errors occurred in both.j2.json confounds[5] · medium/metadata.json world_state
Nine existing E3 owners were switched to medium around 03:04 by an actor outside Fable; effort drift would invalidate a clean comparison.confound; later resolved as Shaan's own change (source/2026-09-06-0345)
Sequential fixed order, cache state, shared-account/server load and small sample may affect timing and usage.confound, unmeasured
The Mini's global config hash changed during the medium arm, by an unobserved actor; both arms ignored user config.confound, unmeasured
The xhigh arm's initial SQLite metadata lookup failed; rollout recovered by thread ID.recovery path, not a result
Limits
  • Effective effort is not independently persisted in either exec rollout; both null values cannot be read as medium, xhigh or equal.
  • First-correct timing is the earliest captured passing version, not an instrumented model intention or server timestamp.
Version
j2.json status measured_with_limits · registered 2026-09-05T20:48:30Z
2026-09-06
06

What it changes

07

Agent entry

00_AGENT_ZERO/domains/compute/runs/2026-09-06-j2/preregistration.json
  1. cd 00_AGENT_ZERO/domains/compute/runs/2026-09-06-j2
  2. jq '.medium.usage, .xhigh.usage' acceptance.json
  3. jq '.medium.records[] | select(.type=="turn_context") | .payload.collaboration_mode.settings.reasoning_effort' runtime-audit.json
  4. node run-pair.mjs # re-run both arms on the Mini (native route)
Agent entry
Read first
00_AGENT_ZERO/domains/compute/runs/2026-09-06-j2/preregistration.json
Owner
GQ-COMPUTE (closed) · last writeback 6 Sep 13:36
Machine
https://siso-shell.pages.dev/t/U19/example.json
Done when
Open this page and see two arms, 122,427 vs 122,327 input, ratio 1.00×, reasoning_effort null in both, and the verdict; then compare against acceptance.json.