Limitations

What we can't claim

Everything else on this site argues for putting a pen in your hand. This page is where that argument runs out, and we would rather you heard it here than worked it out later.

A project that keeps its limits out of sight is asking you to trust it. This one would rather you checked. Every quoted block below is read out of paper-session/references/evidence.md when this site is built, so nothing gets softened on its way to a marketing page. If a limitation here reads gentler than the file, the file is right and the page is broken. The whole brief is on GitHub.

evidence.md · opening

Research support for the paper-session thesis, ordered by the priority ranking. Contested findings are flagged rather than buried, because a public thesis piece that overclaims gets dismantled in the comments, and the honest version of this argument is stronger than the hyped one.

Start here

7 things the research does not show.

evidence.md · What the research does NOT support
  1. No study tests this artifact. There is no research on AI-generated structured worksheets completed offline and returned. Every claim here is assembled from adjacent literatures. Assembling them is a hypothesis, not a finding.
  2. Mind wandering as the mechanism is contested, with a well-powered failed replication (N = 443).
  3. Handwriting's memory advantage is weaker than the popular telling. The encoding and generalization evidence is good; simple recall superiority is not established.
  4. Screen inferiority is a default-mode effect, absent when stakes already force depth. A tablet closes part of the gap.
  5. Most of this is undergraduates in labs. The walking meta-analysis notes most participants were post-secondary students. Ecological validity is a standing limitation the researchers themselves flag.
  6. Nearly all of it is correlational on the AI side. Gerlich's headline finding is a correlation; heavy AI users may differ in ways that predict both.
  7. The "dopamine shortcut" framing is mechanistically sloppy. The dopamine literature shows dopamine allocating effort toward salient rewards, not driving low-effort choices (Michely et al. 2020; Walton & Bouret 2018). The supported mechanisms for AI overuse are effort recalibration, friction reduction, and over-offloading — argue those, not dopamine hits.

Read the first one twice. It is not a footnote to the ones under it. It applies to every rule either skill prints, including the ones that sound most settled.

The three numbers this project leans on

Each of them is true, and each is narrower than it sounds.

  1. d = 0.93

    Walking's effect on divergent thinking, meta-analytic, 23 studies.

  2. 58%

    More ideas produced by employees given unexpected idle time, in a natural experiment at real work.

  3. Confidence held steady

    In the AI-assisted design study, output quality and ideational diversity fell while participants' self-assessed creativity did not move.

Walking

Walk before you think and you come back with more ideas, and with more different kinds of ideas, than you would have sitting still. It is the largest effect in the brief. It is also only an effect on divergent thinking: the same review reports a null on convergent thinking, so walking will not help you choose. Count only the randomized trials and the effect gets smaller, the certainty is moderate rather than high, and item five of the list above names this meta-analysis by name as the undergraduates-in-labs case.

evidence.md · Bonus cluster: the physical case (walking)

Directly relevant since your sessions happen on a patio, and this is the strongest effect size in the brief.

Thabane et al., The impact of walking on creative thinking: A systematic review and meta-analysis (PLOS One, 2026), 23 studies, 1,036 participants: moderate-certainty evidence of a large effect of walking on divergent thinking, d = 0.93, holding at d = 0.82 in randomized trials only, and a null effect on convergent thinking (d = 0.16).

The origin study, Oppezzo and Schwartz's Give your ideas some legs (JEP:LMC, 2014, 514 citations): walking increased creativity for 81% of participants on divergent thinking versus only 23% on convergent, the boost persisted after sitting back down, and walking outdoors produced the most novel and highest-quality analogies.

Two design implications: divergent-only (walking helps you generate, not decide), and the residual boost means a walk before a Deep sheet may be as valuable as the sheet itself.


Unexpected idle time

Employees at a real company had their work cut short by a parts shortage. Left with time on their hands and the problem still live in their heads, they produced 58% more ideas over the following three weeks than the colleagues who worked straight through. It is a natural experiment inside one firm rather than a randomized trial, and the study found no effect at all for interruptions that produced no idle time, and none for planned breaks. It is evidence about being interrupted. Nobody has tested a worksheet.

evidence.md · Cluster 4: The cost of staying on screen

Attention residue. Leroy's Why is it so hard to do my work? (Organizational Behavior and Human Decision Processes, 2009, 289 citations) established that people must stop thinking about one task to perform well on another, and that transitioning away from an unfinished task is difficult and measurably degrades subsequent performance. Her 2018 Organization Science follow-up shows a brief "ready-to-resume" plan mitigates it.

The state of the modern workday. Talypova et al. (2025, ACM Symposium on HCI for Work) recorded 15 knowledge workers in naturalistic one-hour sessions: participants deviated from the main task every 3.5 minutes on average, with nearly 60% of off-task time self-initiated, and average focused attention within the main working sphere down to about 6 minutes, a significant decrease from prior decades.

The most useful single study for your argument. Schweisfurth et al., Unexpected Interruptions, Idle Time, and Creativity (Organization Science, 2023) exploited a supply chain shortage as a natural experiment: employees hit by an unexpected interruption that created idle time produced 58% more ideas in the following three weeks than uninterrupted employees. Crucially, the effect did not appear for intrusions without idle time, and did not appear for planned breaks (where people mentally clock out to leisure goals). The mechanism they land on is attention residue: the unfinished problem keeps working on you only if the interruption leaves your mind on the problem rather than reassigning it.

That is precisely the design of a paper session. Not a vacation, not an intrusion. An interruption that leaves the problem live.

The short-form video base rate. Nguyen et al., Feeds, feelings, and focus (Psychological Bulletin, 2025): systematic review and meta-analysis, 71 studies, 98,299 participants — short-form video use is associated with poorer cognition, with the strongest links for attention and inhibitory control. Chiossi et al. (CHI 2023) found short-form viewing degraded prospective memory via context switching; Luo et al. (Behavioral Sciences, 2025) found reduced cue-based preparation after exposure. These are the attention outcomes; the evidence is thinner for learning depth and reasoning, which is where the honest boundary sits.


Confidence that doesn't move

Teams doing design work with an AI got to their ideas sooner and then stopped looking. They explored less ground, settled earlier, and produced work of lower functional quality and less variety than the teams working without one. Asked afterwards how creative they had been, they rated themselves exactly the same. That is the finding this project is built on, because it says the loss does not announce itself from the inside. It is also the only one of the three that is not a benefit, and it comes from the AI-side material, where item six of the list above applies: that evidence is largely correlational, and people who lean on these tools heavily may differ in ways that produce both halves of the pattern.

evidence.md · Cluster 5: What makes this urgent now

This is the material that makes the piece a 2026 argument rather than a timeless one about notebooks.

AI use correlates with reduced critical thinking. Gerlich, AI Tools in Society (Societies, 2025, 853 citations, n = 666) found a significant negative correlation between frequent AI tool use and critical thinking, mediated by cognitive offloading, with younger participants showing higher dependence and lower scores.

Confidence is the mechanism. Lee et al., The Impact of Generative AI on Critical Thinking (CHI 2025, 498 citations) surveyed 319 knowledge workers across 936 first-hand examples: higher confidence in GenAI predicted less critical thinking; higher self-confidence predicted more. They also document that GenAI shifts the nature of the work toward verification, integration, and task stewardship.

The finding that should open your piece. Tsakalerou et al., AI-assisted design synthesis and human creativity (Frontiers in AI, 2026). AI-assisted teams versus human-only teams on design challenges: AI accelerated idea generation but encouraged premature convergence, narrowed exploration, and compromised functional refinement. Human-only teams engaged in more iterative experimentation and produced higher functional quality and greater ideational diversity. And the kicker: participants' self-perceptions of creativity remained stable across both conditions, meaning the deficit was masked by unchanged confidence.

Pair that with Clinton's calibration finding and you have one argument twice: the screen and the model both degrade the work while leaving your confidence intact. Paper is a calibration device.

And the case for structure specifically. Gerlich, From Offloading to Engagement (Data, 2025, n = 150 across three countries): unguided AI use fostered cognitive offloading without improving reasoning quality, whereas structured prompting significantly reduced offloading and enhanced both critical reasoning and reflective engagement. Vendrell et al. (2026) build a design framework on the same premise, with two principles that read like they were written for this project: preserve cognitive friction, and sequence AI-mediated with AI-free phases.

That last phrase is the academic name for the paper loop.

2026 update: the fatigue mechanism, over-offloading, and the first neural data.

Tian & Zhang, Learners' AI dependence and critical thinking (Acta Psychologica, 2025): AI dependence is associated with lower critical thinking, and the path runs through cognitive fatigue — dependence tires the mind, and the tired mind thinks less critically. AI literacy buffered the effect. This names the mechanism behind Gerlich's correlation above.

Wang, Cognitive offloading through digital tools and its relationship with critical thinking, task persistence, and learning depth (Frontiers in Psychology, 2026): offloading through digital tools is associated with lower critical thinking, task persistence, and learning depth. Honest counterweight, and it matters: in some educational settings the same literature finds offloading supporting self-efficacy and persistence. The variable is not offloading-or-not; it is whether the offloading is designed. Undesigned offloading hollows out; structured offloading can scaffold. That is the entire paper-session thesis in one sentence — and it is why "AI should do the difficult things" is the wrong design brief.

Guo & Ye, Meta-cognitive insights into cognitive offloading (Humanities and Social Sciences Communications, 2026): people frequently over-offload, trusting tools more than is optimal even when incentives discourage it. Metacognition does not self-correct; the tool has to be designed to withhold. This is the empirical backing for the README's "when in doubt, withhold" rule.

Kosmyna et al., Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task (arXiv preprint, June 2025): 54 participants across three groups (LLM, search engine, brain-only), EEG recorded during essay writing. LLM users showed the weakest neural connectivity; when later asked to write without the tool, they showed under-engagement; they reported the lowest ownership of their essays and struggled to quote their own work. Limitations stated plainly: preprint, not peer-reviewed at time of writing; small sample; one task. Suggestive, not settled — but it is the only study in this brief that measures the brain during AI-assisted versus unassisted work, and the ownership finding is the one that should worry you: people don't just think less, they feel less authorship over what remains.

How to hold this: the correlational base keeps growing and the mechanisms are getting named — fatigue, over-offloading, reduced ownership. None of it is causal proof that AI tools make people less capable in general. All of it points the same direction: the default configuration of these tools trains reliance, and reliance is the thing that degrades.


If you want to check the work

The brief has three tiers, and the bottom one cannot promote itself.

Part One argues the thesis cluster by cluster. Each cluster names its sources with citation counts and links, and where the literature disagrees with itself the cluster carries a counterweight paragraph rather than quietly dropping the inconvenient half.

Part Two covers session mechanics: the clusters behind each thing a sheet is allowed to print, format by format, with the contraindications for each one.

Part Three is field reports from real sessions. It sits below everything above it by rule. A report can flag a rule for review; it can never be the only thing holding one up.

Cluster index · 20 clusters

  • Part One

    Cluster 1: Offline incubation and mind-wandering fuel creativity

    The honest counterweight

    The honest counterweight. Murray et al., What are the benefits of mind wandering to creativity? (2021) attempted a conceptual replication of Baird across two studies, N = 443, and found no evidence that mind wandering during incubation facilitates divergent creativity. Du et al. (2025) similarly found mind wandering's effect "limited and context-dependent," with conscious reflection during incubation more beneficial than uncontrolled drifting.

  • Part One

    Cluster 2: Paper beats screens for deliberation

    The honest counterweight

    Boundary condition to disclose: Hoch et al. (2025) found no screen inferiority in high-stakes university e-exams (n = 2,250), where deep processing was already guaranteed by the stakes. Screen inferiority is a default-mode effect, not a law of physics. If the context already forces depth, the medium matters less.

  • Part One

    Cluster 3: Handwriting beats typing for thinking

    The honest counterweight

    The honest counterweight. Richardson et al. (2023) replicated the drawing effect but did not find handwriting superior to typing for free recall, concluding that if the pen is mightier, the effect isn't explained by visual attention or sensorimotor differences. A commentary by Pinet et al. (2025) in Frontiers pushes back on the connectivity paper.

  • Part One

    Cluster 4: The cost of staying on screen

  • Part One

    Cluster 5: What makes this urgent now (the agentic-era case)

  • Part One

    Bonus cluster: the physical case (walking)

  • Part Two

    Cluster 6: Physical manipulation (the correction)

  • Part Two

    Cluster 7: Sketching (the keystone)

  • Part Two

    Cluster 8: Idea selection (now with actual answers)

  • Part Two

    Cluster 9: Analogical distance

  • Part Two

    Cluster 10: Desirable difficulties, and the adoption problem

  • Part Two

    Cluster 11: The return trip (ballot design, forms science, and machine reading)

  • Part Two

    Cluster 12: The after-action review

    The honest counterweight

    The honest counterweights: written debriefing has not been shown superior to verbal (Niu et al. 2021, DOI 10.1016/j.nedt.2021.105113) — the pitch is AAR-versus-no-AAR, with paper's own calibration case (Clinton 2019, Cluster 2) on its separate line. Nearly all evidence is conversational debriefs (~18 minutes average) in training and simulation with performance measured on a repeat of the same task; a one-off solo paper retrospective is an adaptation no study tests.

    Contraindications: never for emotional processing after a distressing or traumatic event — critical-incident stress debriefing was excluded from the meta-analytic base and psychological trauma debriefing has a separate null-to-harm literature; skip when no recorded plan or log exists (a record-free "reflect on your work" sheet exits the evidence entirely); skip trivial sessions.

  • Part Two

    Cluster 13: The premortem (prospective hindsight)

    The honest counterweight

    The honest counterweights: no journal RCT, no study of downstream project outcomes, and no test of a solo pen-and-paper version — the group element (hearing others' reasons) is part of what practitioners credit and is exactly what the solo sheet removes. Gandhi et al. (AoM Proceedings 2022, DOI 10.5465/ambpp.2022.12186abstract; proceedings abstract only, no journal version): premortems mute overconfidence but skew identified risks toward factors outside one's control — hence the "our own doing" input constraint. Ease-of-retrieval: listing 10-12 counterfactuals feels difficult and restores the very confidence it was meant to reduce (Sanna, Schwarz & Stocker 2002, DOI 10.1037/0278-7393.28.3.497; Sanna et al. 2004, DOI 10.1111/j.0956-7976.2004.00704.x) — the 6-8 slot cap is a sourced exception to the overbuild rule, which is grounded in ideation where the goal is novelty; here the goal includes confidence calibration and a demanding quota inverts the mechanism. Field gap the division of labor closes: teams averaged 17.8 failure reasons but few revised plans (Roose & Veinott, HFES 2023, DOI 10.1177/21695067231193680).

    Contraindications: no concrete committed plan — on open-ended ideation the frame is just negativity with no target; an already anxious or pessimistic planner — every supporting study targets overconfidence, and deflating the deflated has no evidential support; repeat runs on an unchanged plan.

  • Part Two

    Cluster 14: The outside view / reference-class forecast

    The honest counterweight

    Contraindications: no defensible reference class (genuinely novel domains — both K&T and Flyvbjerg flag singular cases as the method's limit); unstatable base-rate provenance; an estimate that is not the human's to commit to; any generative sheet.

  • Part Two

    Cluster 15: Teach-back (self-explanation and retrieval)

    The honest counterweight

    The honest counterweights: the novice-audience frame adds no measured advantage in writing and was worse for transfer than plain self-explaining (Lachner et al. 2021) — kept only as a concreteness device, never sold as active ingredient; prompts occasionally backfire (van Peppen 2018) and are stronger with quality scaffolding this sheet deliberately withholds (Rittle-Johnson 2017); closed-book prompts were detrimental after passive processing (Hiller et al. 2020) — hence the engagement trigger; the retrieval advantage reverses at short intervals and will feel worse than rereading precisely while it works (Karpicke & Roediger 2008; Kirk-Johnson 2019, Cluster 10); contested at high element interactivity, negative for relational integration (Mulligan et al. 2023, d = -0.37), and repeatedly fails for problem-solving from worked examples (Huang 2023; Yeo & Fazio 2018). No study tests handwritten teach-back of AI-explained material returned by scan; the composite is a hypothesis assembled from individually well-replicated pieces.

    Contraindications: procedural or worked-example material; no genuine prior engagement with the source; any need for a generative sheet; presenting one sheet as a retention program.

  • Part Two

    Cluster 16: The weekly review (open loops and capture)

    The honest counterweight

    Contraindications: no genuine machine-visible inventory; daily frequency; any quadrant, matrix, or scoring apparatus; fusing with idea-selection or introspection sheets.

  • Part Two

    Cluster 17: The Grinnell field-note system

    The honest counterweight

    Contraindications: Light sessions; generative ideation; merging with reflective formats (excluded by the lineage's own register); projects with no recurring entities.

  • Part Two

    Cluster 18: The serial disclosure kit (expressive writing)

    The honest counterweight

    Contraindications: acute distress, recent or severe trauma, or anyone seeking trauma processing (WET territory — refer to a clinician); self-described low emotional expressiveness; distanced instructions on a disclosure page; any framing as therapy, treatment, or a mental-health intervention; promising relief instead of disclosing the dip.

  • Part Two

    Cluster 19: Brainwriting rounds (group ideation)

    The honest counterweight

    Contraindications: solo sessions, ever; fewer than three or more than six hands; selection anywhere on a passing sheet; printed timers, round promises, idea arithmetic, or memory instructions.

Where the brief argues against its own design

evidence.md · The tension you should address head-on

Sio and Ormerod found that high cognitive demand during the incubation period shrinks the incubation effect. A dense worksheet is a cognitively demanding task. Taken naively, that finding argues against paper-session Deep mode.

The resolution is to stop calling the paper session an incubation period, because it isn't one. In Wallas's terms it is preparation, and Sio and Ormerod found that longer preparation produces a larger subsequent incubation effect. The worksheet loads the problem deliberately and deeply; the incubation happens afterward, in the shower, on the walk, overnight, before you scan.

This reframe also fixes your two-mode design:

  • Deep mode is preparation. Effortful by design, and it earns its cost by making whatever follows more productive.
  • Light mode is closer to true low-demand incubation. Couch, TV, one pen gesture per item. Which is exactly the condition Sio and Ormerod found outperformed rest.

Stated that way, the two modes aren't a UX convenience. They're two different stages of the Wallas model, and the skill happens to implement both.


Taken at face value, that finding argues against Deep mode entirely. The brief prints it in full first and answers it second.

Two paths with nothing under them

Two ways through this loop have no research behind them at all.

If you can't print, there is a card you copy by hand into a notebook you already own. If you can't read a printed page, that same card comes back dictated, typed, or written down by someone else. Both of those leave the literature behind. The brief keeps a heading for each, adds no sources under either, and records cluster by cluster what stops applying. The second path loses more than the first, and the file says so in its opening line rather than somewhere near the end.

evidence.md · The unprinted page (what none of this tests)

Added with the notebook fallback, which serves a user who cannot print by dictating a setup card they copy into a notebook they already own; the same card now also serves a user who cannot read the printed sheet, and what that second trigger costs is recorded in the entry below rather than here. No study tests a dictated scaffold or a hand-copied AI worksheet. The clusters above measure printed pages, and the standing caveat gains a layer here: the artifact is now one the human transcribed before working on it. This entry adds no source. It records what the existing ones do and do not cover once the print step disappears.

"Hand-copying deepens engagement" is a hypothesis, and nothing this skill prints or says may state it as a finding. Cluster 3's encoding half makes it plausible — handwriting produced widespread connectivity in the bands tied to encoding, and handwriting practice generalized to untrained tasks where non-motor practice did not — but none of that evidence measured a person copying a structure someone else authored, and Cluster 3's own counterweight applies unchanged: simple recall superiority is not established. The defensible version is that copying is preparation in the sense Deep mode already claims, loading the problem deliberately. The version that gets sold to a user is neither.

For machine content the fixation side cuts the other way, and that asymmetry is load-bearing. Copying is deeper exposure than glancing. Cluster 9's far-and-uncommon result and Cluster 7's vagueness-fights-fixation mechanism both describe machine-supplied specifics as something to meet briefly and leave; a hand-copied line is the opposite of briefly. Exposure that is harmless where reacting is the task is corrosive where generating is, which is why machine handles cross onto reaction and selection pages only and never onto a page carrying a generative zone. That is two literatures read against each other, not one experiment comparing them.

The grayscale rationale does not transfer to an unprinted page. Cluster 11's dropout-ink inversion — zero-chroma print means any saturated stroke is provably the human's, which is what lets scan-back attribute authorship without registration marks — needs print in order to invert. In a notebook every stroke is the human's, the copied scaffold included, so authorship stops being recoverable from chroma and a printed ink key has nothing to anchor against. That is the whole reason for the dominant-ink rule: hue carries meaning only where two or more inks appear, and a page written in one ink is read by imperative form — also the safer reading, given the same cluster's finding that blue and black are the least separable pair in a phone photo.

Expect the copy to feel like wasted effort. Cluster 10's adoption finding — effortful strategies are judged less effective and chosen less often, and are the ones that work — predicts transcription as the loop's most likely abandonment point, and predicts that the people who finish it will rate it poorly. Whether anyone copies a card and returns the pages is unmeasured. That is a completion question rather than a lab question, and Part Three is where its answer lands.

evidence.md · The unread page (what none of this tests, and what it costs)

Added with the non-visual return, which lets a person who cannot read the printed sheet take the same setup card and come back dictated, typed, or written by a scribe. No study tests a dictated, typed, or scribed return of a generated worksheet. This entry adds no source, because there is none to add. It records which of the clusters above stop applying when nobody looks at a page, cluster by cluster and by name, because the loss on this path is larger than on any other in this file and understating it is the dishonest way to ship an accommodation.

The trade, stated before the ledger. This path gives up most of the mechanisms the clusters above measure, and it gives them up to reach people the loop currently cannot serve at all. It is the right trade, and the reason is that the alternative is not a thinner loop for them but no loop. It is not an equivalence, and nothing either skill prints or says may present it as one — not as "the important part survives," not as "mostly the same benefit." What survives is named below, and it is the smaller half.

Lost wholesale: Cluster 2, paper beats screens for deliberation. The cluster is a paper-versus-screen contrast end to end, and a spoken answer is neither medium. The g = -0.25 comprehension gap and the g = 0.20 calibration advantage this file calls its most important number are both unavailable here. One part of the cluster still speaks, and it speaks against the typed variant rather than for it: Sidi et al. found screen inferiority on tasks with almost no reading burden at all and concluded the medium itself is a contextual cue licensing shallower processing. A return typed on a screen inherits that cue by the same reasoning. Predicted, not measured — nobody has run it.

Lost wholesale, and lost by name: Cluster 3, handwriting beats typing for thinking. A typed return is literally the cluster's control condition; a dictated one is outside the comparison altogether. A scribed return needs its own sentence: whatever the motor act contributes, on a scribed page it is the scribe's hand performing it, so the encoding evidence — weak already, since Cluster 3's own counterweight is that simple recall superiority is not established — accrues to the wrong person.

The largest single cost: Cluster 7, sketching. This file nominates sketching as the strongest paper-beats-screen mechanism in the entire brief, and it is the one lost with no substitute. Stones and Cassidy's comparative result and Tversky and Suwa's vagueness-fights-fixation mechanism both require marks a person makes and then looks at again; the whole finding is reinterpretation of one's own ambiguous strokes. Nothing in this file substitutes for that, and inventing a stand-in activity would be a printed rule with no source. This is also the only path on which the Deep dot-grid requirement yields, which is why it is scoped in writing rather than dropped by inference: where the sketch cannot be worked, the kit carries the prompt the drawing was there to ask, and nothing takes the drawing's place.

Lost wherever it was earned: Cluster 6, physical manipulation. Cut-apart cards go with the page. The cluster's own correction already scoped them to tasks with many items that are relational or spatial and exceed working memory, so what is lost is narrow — but on this path it is total, since arranging items in space is the mechanism.

Reduced to a fraction: Cluster 11, the return trip. Four of its parts are void here. The mark vocabulary — strike, circle, ?, !, star, arrow — is visual-spatial. The dropout-ink inversion that lets scan-back attribute a stroke by chroma needs zero-chroma print to invert against, and is void for the same reason it is already void on an unprinted page. The layout-is-instruction half — box size driving word and theme count, multiple small numbered slots outperforming one open box, per-row forced choice beating check-all — is a set of findings about answer-space geometry, and there is no answer space. So is the marking-gesture half, where filling an enclosed target roughly halved lost votes. What survives is the part the cluster names as primary anyway: wording. A spoken "cut that" is a strike. One thing runs the other way and should be said plainly rather than buried: a typed return skips handwriting transcription entirely, so the cluster's named weak spot — fluent plausible-word substitutions, worst on blurred and rotated captures, with struck-through text the most machine-expensive mark — does not arise. A dictated return moves that risk to speech recognition, which no source here measures. A scribed return keeps handwriting transcription and adds the scribe's own copying error underneath it.

Cuts against this path rather than being lost by it: Cluster 10, desirable difficulties and the adoption problem. Dictation is the least effortful of the three return channels, and Kirk-Johnson's finding is that the effortful strategies are the ones that work and the ones people decline. Expect a dictated return to be the shallowest of the three, and expect that not to be how it feels.

Survives, and cuts both ways: Cluster 4, the cost of staying on screen. Schweisfurth's result is about an interruption that creates idle time and leaves the problem live, which a handed-over card does as well as a printed sheet does. Leroy's attention-residue half is the caution on the same page: a dictated return lands back on-screen, in the chat, which is the surface the interruption was meant to leave. The thinking is off-screen at both ends; the handover is not.

Survives: Cluster 1 and the walking bonus cluster. The off-screen half of the thesis does not require paper. Sio and Ormerod's second finding — longer preparation produces a greater incubation effect — is served by a dictated card as well as by a sheet, because the card is the preparation. Walking's d = 0.93 on divergent thinking is the largest effect in this file and pairs with dictation rather than against it. This is the honest reason a non-visual session is still this loop and not just a chat: what is lost is the paper mechanisms, not the leaving.

Survives because it was never about the medium: Clusters 5, 8, 9, and 12 through 18 — with three qualifications, none of them a new rule. Cluster 8's answers are instructions rather than geometry, and its null on ranking mechanics means a dictated ordering is as good as a marked one. Cluster 14's distribution strip is a spatial self-placement gesture with no tested dictated form, though the corrective Buehler's Study 4 actually measured — apply your typical past completion time and write a scenario in which the current case goes typically — is verbal and carries over intact. Cluster 18 is the sharp one: its dose is already this skill's untested conversion from tested minutes to one full page, a dictated version converts it a second time, and nothing in this entry licenses returning a serial disclosure page by any channel. That page is defined by never coming back, the marker printed on it says so exactly, and a dictated return is a return.

Has no non-visual form this file can support: Cluster 19, brainwriting rounds. The passing is the mechanism — Paulus and Yang's winning condition was reading the ideas already on the incoming sheet and adding to them, and the sketch-permitting replications beat word-only cells — and reading the incoming sheet is precisely what this path cannot do. No source here says how a non-visual participant joins that table, and this entry does not invent one.

What this file still has no source on at all. Screen readers, PDF semantics and tagging, braille, and motor impairment. Cluster 11 carries the only two passages in the file that read as accessibility work — the color-vision-deficiency rule (Jeong et al. 2025, roughly 1 in 12 men) and the psychophysical legibility floors (Legge & Bigelow 2011; Chung et al. 1998; Akutsu et al. 1991; Crossland et al. 2012's 29.6% of 65-84-year-olds) — and both describe a person reading ink on paper with their eyes. Neither says anything about a person who does not read the page, and neither may be cited as if it did. The document metadata the generator sets — /Lang, /Title, and DisplayDocTitle — rests on standards conformance (WCAG 3.1.1, WCAG 2.4.2, and PDF/UA-1's requirement that a document title be accompanied by DisplayDocTitle) rather than on a measured effect, and Cluster 11's own standing for citations of that kind is inherited here unchanged: the standards encode incident experience, not effect sizes. Whether a real reader announces any of it is untested, and whether an untagged sheet is navigable at all is not in question — it is not.

Whether anyone takes this path is unmeasured, as is whether the people it is for would rather use a scribe, a braille slate, or a typed file, and whether the setup card survives being heard rather than read. Those are completion questions rather than lab questions, and Part Three is where their answers land.

The empty tier

Nobody has reported finishing a session.

Part Three is where real sessions land: date, capacity, patterns used, whether the pages came back completed, partial or abandoned, and where they stopped. It has no entries. Whether people finish is the only measure this project treats as success, and there is no data on it yet. That is also why you will find no testimonials on this site and no percentage anywhere claiming that people finish.

evidence.md · Part Three: Field reports

The tier below everything above. Real sessions, reported by the people who ran them. This section exists so field data has a defined place to land — and a defined ceiling.

Rules for this tier:

  • Every entry states its N and how it was collected. Most entries will be N = 1 and self-selected; say so.
  • Field reports sit below the lab evidence, always. A report can flag a rule for review; it can never be the sole support for adding one. The bar in CONTRIBUTING.md still holds: a rule that changes what gets printed needs a source in the clusters above.
  • Limitations up front, in the same register as "What the research does NOT support."
  • What an entry records: date, capacity (Deep or Light), patterns used, what came back — completed, partial, or abandoned, and where it stopped — plus anything the scan-back pass misread.

Threshold for action: three independent reports agreeing on the same failure — the same pattern abandoned, the same zone misread, the same sheet type not coming back — justify opening a pattern-review issue. Until then, entries accumulate here.

No entries yet. Completion rate is the project's missing metric; this is where it stops being missing.

What is verified, and what only installs

Installing it is not the same as having run it.

Rows in the compatibility ledger that report the loop tested from end to end: 1. Every other row says something weaker, either supported or installs cleanly with the loop untested, and nothing on this site rounds any of them up.

Where the loop stands, agent by agent

Your agentInstallThe loop todayAlso install
Claude Code (CLI, IDE, web)npx skills add, or clone-and-unzipVerified end to end — discovery, install, sheet generation, and scan-back all testedreportlab + pdfplumber, on the machine running the agent
claude.ai · Claude desktop · mobileupload the two .skill filesSupported — the chat-native path; a phone photographing the pages is what the return trip was built aroundNothing by hand — the skill installs what it needs where it runs
Cursor · Codex · Copilot · other CLI-routable agentsnpx skills add … -a <agent>Installs cleanly, loop untested — the sheets need an agent that can run reportlab and read photographs back. A session report from one of these is the contribution we want mostreportlab + pdfplumber, wherever that agent runs Python
Anything else that reads SKILL.md foldersdownload a .skill (it's a zip), unzip into the agent's skills directorySame as aboveSame as above
Chat AI with no skills runtime — ChatGPT, Gemini, Copilot chat, Perplexity, Le Chat, DeepSeek, Grok, Poepaste paper-session-paste.md, then scan-back/SKILL.mdUntested, per surface — the forward half produces a hand-copied card, never a PDF, and nothing is verified because nothing is rendered. A pasted protocol also has no authority over the host's system prompt, so the bright lines hold only as far as the host chooses to hold themNothing — no PDF is generated, so no library is needed

The same table, and the install directions it belongs to, are on the install page.

None of this gets settled by reading more of it. It gets settled by pages coming back with handwriting on them.

Get a sheet

Open territory