~/runthismodel
daemon okbuild 5a3c91d00:00:00Z
verified runs / advanced-creative-continuity-v1
Creative continuity·protocol 1.0.0 · 2026-08-03Measured

Qwen3.6 35B — Long-Form Creative Continuity

Qwen3.6 35B wrote long-form prose locally on Ollama for Windows. Fourteen deterministic checks covered headings, chronology, object state, event order, exact text, dialogue count, and length. Every attempt and each check is shown as measured.

model / workflow
qwen3.6:35b
task
advanced-creative-continuity-v1
checks held
11 / 14
measured attempts
4
best rate / timing
73.2 tok/s
attempt ledger·every attempt shown
attempttotalloadratecheckswhat happened
Cold initial51.83 s24.41 s71.75 tok/s10/141,503 words; four contract failures.
Warm identical27.66 s0.29 s71.22 tok/s10/14Byte-for-byte identical to the cold response.
Warm repair 128.71 s0.29 s73.22 tok/s11/14Best run; missed by two words, one time, and one verb.
Clean repair 221.31 s0.32 s73.06 tok/s10/14Narrower feedback regressed to 1,146 words.
What the checks cover

The first draft created a flooding lunar railway station, carried a sealed envelope through rising water, and gave two characters recognizable voices. It read like a real story.

These fourteen checks measure explicit instruction control, not literary merit: headings, chronology, object state, event order, exact text, dialogue count, and length. Voice, imagery, causality, subtext, rhythm, and the ending sit on a separate blinded editorial worksheet.

How the runs measured

The first targeted repair reached 11 of 14. It ran 1,602 words — two over the ceiling — omitted one required timestamp, and changed an exact required verb.

A narrower second repair returned to 10 of 14 at 1,146 words. One additional run is kept but excluded from the speed comparison because a supervised ComfyUI process restarted during the protocol and occupied the GPU.

What this shows

The prose was usable as a draft. Long-form fiction, game narrative, screenplays, and branded content all carry state that polished language can quietly contradict, so the explicit contracts are checked directly.

  • The model is used for invention.
  • Deterministic checks cover the explicit contracts.
  • A human editor covers literary judgment.
  • Excluded runs are kept, with the exclusion reason stated.
reproduce / inspect·downloadable evidence
Measured results
Four clean runs, one excluded run, word counts, timing, and scores.
Validator source
Fourteen deterministic continuity and instruction checks.
Scope boundary
The percentage measures explicit instruction control, not literary quality. It is not a universal model ranking or hardware purchasing recommendation.