By the end of one day at NExT I had held an entire client system in my head well enough to answer questions about it on demand, planned what the next engagement would need, designed schemas that other people's work would be built on top of, watched several Claude Code sessions against a token budget tight enough that a wasted run actually cost something, and been present — genuinely present, not nodding — through a run of design conversations.
Then I went home and played Silksong for two hours.
The obvious explanation is that I'd used myself up. Eight hours of concentration, tank empty, collapse onto the couch. It's the story everyone tells, it has a satisfying shape, and I'd have told it myself.
It's wrong, and the tell is the game.
Silksong is not a couch. It's a precision platformer that kills you for a late input and then makes you walk back. If I had genuinely run out of whatever concentration is made of, I would have watched something. Instead I spent two hours on a task demanding exactly the fast, sustained, error-punishing attention I had supposedly just exhausted — and I was happy to be there, and I was good at it, and I slept fine and woke up fine with no sense of having borrowed anything or paid anything back.
So what had I actually run out of?
Nothing. Nothing had run out. Something had been repriced.
I know how that sounds. It sounds like a nicer way of saying I was tired. It isn't, and the gap between those two sentences costs about 20% more work per week — I'll put that number on the table later. It's also the gap between two entirely different models of what software work costs. Every tool we use to plan that work is built on the model that's wrong.
I'd know. I wrote an essay on it.
I got the unit wrong
In May I argued that good code is code that saves you time →. Readability, performance, modularity, simplicity — none of them sacred, all of them proxies, each earning its place because somewhere down the line it saves somebody time. Arguments about style, I said, are really arguments about whose time gets spent.
I still think the structure of that argument is right. Every quality property is downstream of some resource, and the fights are about who pays.
I got the resource wrong.
Time is measurable, not scarce. Those are different properties, and we have confused them so completely that the entire apparatus we use to plan software — estimates, story points, sprint capacity, velocity, the calendar itself — is denominated in a unit that does not bind. It's the streetlight problem, running for decades at industry scale. We count hours because hours hold still while you count them.
Herbert Simon named the actual constraint in 1971, before most of this apparatus existed:
What information consumes is rather obvious: it consumes the attention of its recipients. Hence a wealth of information creates a poverty of attention and a need to allocate that attention efficiently among the overabundance of information sources that might consume it.
Fifty-five years. Every industry that sells eyeballs took this seriously and built an economy on it. The industry that consumes attention to produce rather than to sell has, as far as I can tell, never once changed a planning artifact because of it.
The claim in this essay is that attention is the binding resource in software work, and that it differs from time in five specific ways — each of which costs something concrete, and none of which any common estimation system can represent. Then a tool I built that my own benchmark said not to build, and what to do about all of it.
A warning about the middle section before we get there. I'm going to tell you what attention is, using research, and I want to be honest that the research is a live fight rather than a settled foundation. It matters less than it sounds like it should. You can reject every mechanism in the next section and the phenomena they were proposed to explain are still sitting there — maybe caused by something else, maybe described incompletely, but observable without a lab and large enough to plan around. The argument runs on the phenomena, not the explanations.
What attention actually is
Start with the concession, because it's load-bearing.
"Attention is a resource" is a metaphor, and a contested one. Simon's version assumes a central allocator handing out a unified, general-purpose supply, and both halves of that are rejected by a good deal of current work — the philosopher Jelle Bruineberg argued last year that the term may not even have a stable referent, which would leave the whole attention economy without a coherent foundation.1
Fine. You don't have to believe a theory of mind to notice five properties that show up in your own workweek whether or not the neuroscience ever settles, because I can point at each one in a repository.
The companion to this essay → states four of these in a paragraph and gets on with its argument. This is the long version, and the long version is where the useful parts are.
It's a flow, not a stock
Time is a stock. You get a fixed allotment, you spend it, unspent time is simply later.
Attention doesn't work like that. Unused attention evaporates. There's no account it accrues in, no way to be extra attentive tomorrow because you were bored today. Attention economists classify it as a flow currency for exactly this reason: it cannot be saved for later, and it's forgone if it isn't spent.2
Which means idle attention is destroyed value.
That reframes boredom as a loss rather than a rest, and it also explains a thing I've never seen a productivity system account for: you can be extremely busy and extremely bored at the same time. Boredom isn't idleness. The cleanest definition I know comes from a 2012 paper defining it strictly in terms of attention: boredom is the aversive state of wanting, but being unable, to engage.3 Having attention with nowhere it can go.
I have spent whole afternoons fully occupied and thoroughly bored. So have you.
It isn't one pool
The intuition that you have "an amount of focus" and tasks draw it down is the single-pool model, and it was replaced decades ago in the field that actually has to predict this stuff for a living — human factors, where getting it wrong crashes aircraft.
The successor is multiple resource theory. Two tasks interfere with each other in proportion to the machinery they share: perceptual versus response stages, visual versus auditory input, spatial versus verbal encoding.4 Different channels, cheap concurrency. Same channel, brutal.
The engineering consequence is that what you can run in parallel is set by overlap. Reviewing a diff while a build runs is nearly free. Two design conversations back to back are not, and neither is code review sandwiched between them, because all three are the same faculty wearing different clothes.
This is also why "I have four hours open" tells you almost nothing about what you can put in them.
It bills at different rates
Some work gets better when you concentrate harder. Some work does not, and the distinction is old and precise: Norman and Bobrow separated resource-limited processes, where performance improves as you allocate more capacity, from data-limited ones, where performance is capped by the quality of the input and additional effort returns exactly zero.5
Zero. Not diminishing. None.
When I was a TA we had pylint wired into the autograder, and it produced somewhere between one and two hundred lines of complaint per submission. Twenty TAs, twenty submissions each, and the job was to find the handful of infractions that mattered inside all that noise. Everyone's instinct — mine included — was to concentrate harder and scroll more carefully.
That instinct was worthless, because the task was data-limited. The information was all present and arranged badly, and no amount of additional attention rearranges it. So I wrote something that grouped the output by file, then by category, then collapsed the rest to line numbers, and the task stopped being hard, because the data changed shape →.
Half of what passes for engineering discipline is people focusing harder on output that cannot reward focus.
Open work encumbers it
This is the property with the best evidence behind it and the least uptake anywhere.
When you stop working on something before it's finished, part of you keeps working on it. Sophie Leroy named this attention residue — the persistence of cognitive activity about task A while you are ostensibly performing task B — and the finding that matters is that incompleteness is the strongest predictor of how much residue you carry.6 The unfinished thing keeps billing you after you've left the room.
The word for that is lien — an encumbrance on future attention you didn't choose to take on, that doesn't clear by working harder, and that keeps charging until it's discharged.
And here's the part nobody ships: you don't have to finish the task to discharge it. Leroy's follow-up work, and an independent line by Masicampo and Baumeister, both found that writing down a specific plan for the unfinished thing eliminates the intrusion and the performance cost — while the goal is still open.7
The rent is charged by unresolved state, not by an open task. Externalize the state and the lien clears.
A TODO with a concrete next action costs nothing. The same TODO in your head costs rent until it closes.
Sustained control reprices it
Which brings me back to Silksong.
The leading account of what happens after prolonged effort is a process model: exerting control causes temporary shifts in motivation, away from what you have to do and toward what you want to do, and in attention, away from the cues that signal a need for control and toward the ones that signal reward.8 Nothing is consumed. What changes is what you're willing to spend on.
That's the evening, exactly. I hadn't run out of the capacity to concentrate — I demonstrated that for two hours straight against a game that punishes lapses. What I'd run out of was any willingness to spend concentration on work-shaped problems. The price had moved, and the price had moved because I'd spent eight hours doing nothing but override my own defaults.9
The same mechanism has a much less charming form at work, and I have receipts for that one too.
At my worst — five lanes, three of them implementation and the others capturing screenshots and waiting on a review, with no tooling holding any of it for me — I didn't get worse at engineering. I got selective about which steps I paid for. I'd smoke-test only the surface I'd just told the agent to fix — a seemingly safe shortcut, and one that was wrong often enough to be predictable, because things underneath it broke constantly. I'd check the screenshot instead of the live site. The fixtures are accurate; that one random hover state still dies inside.
Both of those are steps whose omission is invisible at the moment you omit it. That's the pattern. Rationing targets the unwitnessed. My review quality never slipped — I don't think I ever approved someone else's work that needed a rollback — while reviewers were catching genuinely amateur mistakes in my own branches. A review is externally scheduled, bounded, and visibly attributable to me. Skipping a smoke test is none of those things.
Not all of those mistakes got caught. A few reached staging and needed a rollback.
Any individual shortcut was defensible. That's what makes it dangerous, and it's the same arithmetic that justifies automating in the first place, pointed the other way: trivial to not make the mistake once, a statistical certainty across a thousand chances.
Twenty-five years of the wrong physics
There's a reason I keep saying repriced instead of drained, and it's worth a short detour into a field that spent a quarter century making precisely the mistake I'm accusing our industry of making.
Psychology had a beautiful theory called ego depletion: self-control is a limited resource, like a muscle that fatigues. It gave us decision fatigue, the glucose hypothesis, and most of the willpower canon, and a 2010 meta-analysis across 198 tests put the effect at a very healthy .
Then people tried to replicate it under preregistration, where you commit to the analysis before you see the data.
The 2016 multi-lab replication — 23 labs, — found , with a confidence interval that comfortably contained zero. The theory's own authors objected that the chosen task was a poor test — fairly or not, one of them had approved it beforehand — so a 2021 effort crowdsourced experts to pick tasks judged representative of the phenomenon instead: thirty-six labs, , , with the Bayesian evidence favoring no effect at all. Then in 2025 a third group argued the nulls came from manipulations too weak and too short to deplete anything, ran a paradigm with real teeth — a demanding antisaccade task lasting thirty to forty minutes before the measurement — and got it back. Fourteen samples, , to , with essentially zero heterogeneity between sites.10
So the phenomenon exists. Push someone hard enough for long enough and their subsequent self-control measurably degrades.
But look at what the surviving paradigm requires. Thirty to forty minutes of sustained demand. That's a time-on-task effect. In twenty-five years and thousands of studies, nobody has ever demonstrated the tank — and the mechanism that replaced it, the one in the previous section, describes shifts in motivation and attention instead of consumption. The theory's own originator now describes the modern version as emphasizing conservation rather than exhaustion.11 You don't run out. You start rationing.
They were doing economics the entire time and calling it thermodynamics.
I want to be careful about how much I'm claiming here, because it would be cheap to say the productivity canon is built on a corpse. It isn't. The advice mostly works. What was wrong was the mechanism, and getting the mechanism wrong is expensive in a specific way: it makes you prescribe rest when the problem is price.
I made the identical error myself, one level down. I described my own bad days as going "on cooldown" — a tank refilling, wait it out. The folk model predicted the right behavior, which is why it survived. It predicted the wrong intervention every single time.
Software's version of this error is one level up again, and much older. We call attention time, because time is what we can put on a board.
A tool I'm faster than
During a styling migration I built a Playwright screenshot harness: a deterministic capture of a large set of surfaces against pinned baselines, because the build stayed green while the interface broke. The companion has the full account →; that sentence is all you need here.
Then, later, I benchmarked it against myself. Five focused runs, honest conditions.
I'm faster than it is. Comfortably. I can leave the pages I need already open, and my incremental recompile trounces the harness's compile-from-scratch on every single capture cycle. At full focus I also catch more than it does, because a real browser in front of me has every state in it and the fixture set only has the ones somebody thought to build.
By every measurement anyone would actually take, the tool loses to the human. The correct conclusion from that benchmark is don't build this.
The benchmark is measuring the wrong integral.
I can beat the harness for five runs. I cannot beat it for four hundred, because I can't stay at that level — the honest ceiling is more like twenty minutes at a stretch, and after about an hour of that I need to go do something else. What actually happens over a day is the previous section: I start skipping the steps whose omission I won't see. Then those become defects, and defects become follow-up pull requests on later days — call it 20% more work per week, on top of the breaks I "had" to take to be able to work at all.
So the harness did save time. Enormous amounts of it — by preventing the rationing that produces rework. And there is no per-run benchmark in existence that can see this, which is precisely why my per-run benchmark got the answer backwards.
Automation's product here is invariance. Code doesn't get bored. Code doesn't decide it'll just check the surface it was told about. My quality has a decay term and the machine's doesn't, so the integral over a workday inverts every point comparison — and the point comparison is the only one anybody runs.
That sharpens something in the companion essay →. Its second term prices both directions of being wrong — what you pay when the tool does too much, and what you pay when it misses. It's the right question. It just prices both halves against a human who holds still.
I don't hold still. My miss rate at run four hundred is not my miss rate at run five, and the harness's is the same number at both. So the comparison isn't between two error profiles. It's between one that's flat and one that slopes, and which way it comes out depends entirely on where along the span you take the measurement. Ask at five runs and I win in both directions. Ask at four hundred and I lose in both.12
The missing variable isn't inside either error term. It's the length of the span you're pricing them over — and a benchmark, by construction, sets that length to one.
How to spend it
None of what follows is generic productivity advice. You already know about focus blocks and turning off notifications; that advice is fine and it's the hundred level. What follows are consequences of the five properties, and most of them are strange enough that I've never seen them written down.
Experiment far past the point you can justify
The old cost of an experiment was your attention across the entire loop: form the idea, build it, run it, judge it, write down what happened. Five demands, four of them on the expensive channels.
Delegation collapses that to two, and the cheaper one is nearly free. Devising an experiment doesn't draw on the same stream as executing work — it's the difference between a build wait and a second design conversation — so you can design one while something else is running and barely feel it. You can also hand that step over too, ask for five things worth trying, and pick by feel.
What comes back is a single bounded act of judgment, plus, if the harness is any good, a written account of what was tried and what happened. That artifact is precisely the thing that discharges an open loop. An experiment you abandon costs you almost nothing, because you aren't carrying it around afterward.
And the quality bar can fall very far, which is the part people miss. An experiment is the same transaction as an intern saying they'd like to try something: the cost is bounded, the output is information rather than product, and nobody expects it to ship. Code that would never survive review is entirely adequate for telling you whether an approach is alive.
So the calculus inverts. It used to be cheaper to think hard about whether something would work than to go and find out. Now finding out is cheaper than thinking, and the correct number of experiments is far higher than your instincts allow.
One condition, and it's the entire condition: this holds only while your evaluation stays cheap and trustworthy. Cheap generation with expensive evaluation is worse than not experimenting at all, because you have moved the whole cost onto the channel that was already scarce. If you can't tell good from bad at a glance, build that first and experiment second.
Spend boredom on diagnosing the boredom
Boredom is attention with nowhere to go, which means that when you're bored you have surplus capacity by definition — and it's the one resource you can't put away for later. Spend it on the boredom itself.
Boring is a measurement. It has roughly three causes and they take different fixes.
Data-limited: you're concentrating on something that can't reward concentration. Rearrange the information. The TA aggregator → came from exactly this — scrolling lint output was boring, boring meant the information was shaped wrong, and the fix was grouping it.
Latency-bound: you're waiting. The loop is too slow to hold you, and the answer is either a faster loop or a second lane to put the spare attention into.
Genuinely mechanical: there's no judgment in the work at all. That's an automation candidate and you already know the arithmetic.
The mistake is treating boredom as something to endure, or as a cue to go somewhere else. It's free diagnostic capacity, arriving at the exact moment you have something worth diagnosing.
Know when not to open another worktree
Only one in ten interrupted programming sessions resume in under a minute, and the way people get back in is by running the program and re-reading code — rebuilding externally what evaporated internally.13 So the discipline you need is a function of how little of you each lane is going to get. One deep task all day, and your hygiene can be as loose as you like; nobody needs a pinned baseline to remember what they were doing ten seconds ago. Eight lanes at five minutes an hour each, and every one needs a pinned baseline, a scoped diff, and a written next action, because none of those five minutes can go to recall.
But hygiene only buys you so much, and the harder skill is knowing when to stop opening lanes at all. I use two signals, and either one alone means I'm over:
I can't say what each worktree I care about is currently thinking. Not what it was assigned — what it's doing right now, and whether that's still the right thing.
I can't keep up with them coming back to me.
When either fires there are exactly two moves. Devise longer and more independent tasks, and accept the extra risk that buys — a lane you check less often is a lane that can go further wrong before you notice. Or shed load outright and clear the liens.
The number of worktrees you can respond to concurrently is your ceiling, and it always will be. It doesn't move because the agents got faster; it moves only if each lane needs you less often. Project managers worked this out a long time ago and gave it a name — span of control — and then watched people ignore it for several decades.
Price automation in attention
The obvious case is real and nothing here disputes it. Automate something that takes an hour a day and you have saved an hour a day — and with it the attention that hour would have consumed, plus the lien it was holding the whole time it sat on your list. Time savings are attention savings. They're simply the ones everybody already knows how to count, which is why they're rarely where the decision is hard.
What gets mispriced is everything where the clock reports that nothing happened.
At one job I built LunchBot, which is exactly as serious as it sounds: a Slack relay that forwarded the announcement that catered lunch had arrived into an opt-in channel. Before it existed, the way you found out lunch had arrived was to go check the original channel yourself, and people checked something like twenty times a day.
Measure that in time and it's nothing. Two seconds a check, under a minute a day, not worth writing software for. Measure it in attention and it's twenty context switches, each one a small lien opened on arrival and closed on checking, spread across exactly the stretch of the day when you're least able to absorb them. The bot didn't save a minute. It converted polling into an interrupt, and the whole cost was in the polling.
That's the general shape, and it's the most reliably mispriced thing in engineering. Anything you check is costing you far more than the clock says, because the cost is in the checking rather than the thing checked. CI dashboards. A long build. An agent you keep glancing at. Every one of those is you running a poll loop on wetware.
The second class is what you stop doing properly. Ask what you skip at hour six — not what you do most often, but what you quietly stop doing well near your ceiling, and specifically the steps whose omission you won't notice at the time. For me that was smoking every surface instead of the one I had just changed, and checking the live site instead of trusting the fixture. Code has no hour six. The repetitive-but-reliable things can wait; you'll keep doing those correctly.
Batch by faculty, and know what you're trading
Interference tracks shared machinery, so the grouping that matters is by what a task uses, not by what it's about or how urgent it is. Two design conversations back to back look like sensible batching and are close to the worst pairing available. A build wait and a code review look unrelated and cost almost nothing together.
Meetings are where the usual advice gets this wrong, because "batch them" is only half a rule. Stacking meetings does buy you a contiguous block afterward, and it collapses four resumptions into one, which is the thing everyone optimizes for.14 It also concentrates every boring meeting you have into one unbroken stretch of attention with nowhere to go — and that's destroyed value, not rest. You come out of three dull meetings in a row worse off than you went in.
So the trade is real and it depends on the meetings. Ones you're genuinely in — where you're thinking, arguing, deciding — batch freely, because they're consuming attention productively. Ones you merely attend either need to be spread out so the attention has work on either side of them, or batched deliberately with something else to put that attention into. Which is a sentence I'm aware reads as a recommendation to do other work during meetings, and I'm not going to pretend otherwise.
Match arrival rate to the task, and alternate
How loaded you feel is governed by how often work comes back to you rather than how much of your time it takes. Early on, five lanes at 50% utilization had me drowning. A job later, once the delivery loop was encoded and I wasn't the one holding it open, eight lanes at 67% were comfortable, because the interruptions had fallen from sixty an hour to eight.15 More lanes, more utilization, less load. That's a real effect and it makes a strong case for the tool that interrupts you once with everything instead of ten times with fragments.
But minimizing interruptions is the wrong objective. Arrival rate is something you allocate, not something you reduce.
On a task that matters — ambiguous requirements, a migration that can eat a weekend, anything where being wrong is expensive — you want it coming back more often, not less. Models lose the thread. You can watch the reasoning drift and catch it before it's produced anything, which is enormously cheaper than catching it in the diff. Short leash, frequent returns, and you spend the attention on purpose.
On a well-specified task with a cheap failure mode, take the long leash and spend nothing until it's done.
The failure I see most often in people working this way is not picking wrong on any individual task. It's not picking at all: going fully hands-off on everything, or fully hands-on on everything, and treating it as a personality rather than a per-task decision. The skill is alternating, and the tell that you have it is that your leash length varies across the lanes you have open right now.
What's left to be good at
Generation is getting cheap and it'll keep getting cheaper. What doesn't get cheaper is knowing which of the things you could build is worth building, and being able to point a finite amount of yourself at the part that matters.
That's most of the job now, and I think it's nearly all of it soon. The developers who matter will be measured on their taste and their attention — what they choose, and where they spend themselves. Neither has ever appeared on a dashboard.
What this doesn't prove
The evening with Silksong is one evening, recalled, with no measurement of anything. It's an illustration of a mechanism established elsewhere, not evidence for it.
The lane counts, the utilization figures, and the 20% are my recollections and estimates from the period, not instrumentation. I'd defend their shape and I wouldn't defend a decimal place.
The research is ongoing and I've quoted a fight rather than a consensus. The strongest claims here — residue, the data-limited distinction, the resumption times — are the ones with the least controversy behind them. The weakest is the framing itself, which is why I put Bruineberg's objection at the top of the section instead of burying it.
And I'd extend to myself the same courtesy I extended to twenty-five years of psychologists. My own "cooldown" model survived years of daily use because it worked well enough to act on, and it was still wrong. This essay's model may be in the same position. If it is, the part that will turn out to be wrong is the explanation, not the observation.
The observation is that I gave a benchmark an honest question and it gave me a confident, well-measured, wrong answer, because it was denominated in the resource I could count instead of the one I was actually spending.
That part I'm sure about.
Notes
-
Jelle Bruineberg, "Rethinking the cognitive foundations of the attention economy," Philosophical Psychology (2025). The argument is that individual attention emerges from competition among concurrent processes rather than from any central allocation of effort — which is a problem for Simon's framing specifically, since Simon needs an allocator. ↩
-
Maxi Heitmayer, "The Second Wave of Attention Economics: Attention as a Universal Symbolic Currency on Social Media and beyond," Interacting with Computers 37(1), 2025. The stock/flow distinction is his, and so is the phrasing — attention "cannot be saved for later and it is foregone if it is not spent." Note the asymmetry he draws out: the flow can be converted into a stock, as reputation or audience, which is the mechanism the entire attention economy runs on and is not a move available to you on a Tuesday afternoon. ↩
-
Eastwood, Frischen, Fenske and Smilek, "The Unengaged Mind: Defining Boredom in Terms of Attention," Perspectives on Psychological Science (2012). Formally: unable to engage attention with the information a satisfying activity requires, focused on that failure, and attributing it to the environment. ↩
-
Christopher Wickens' multiple resource theory, the standard model in human factors. Its axes are stages, sensory modalities, processing codes, and visual channels. I'm extending it well past its warrant when I apply it to kinds of work rather than sensory channels — reviewing and designing aren't different modalities in Wickens' sense. The extension matches my experience; it isn't what the model says. ↩
-
Norman and Bobrow, "On data-limited and resource-limited processes," Cognitive Psychology 7(1), 1975. Usually cited as the origin of resource allocation theory. The same paper notes that interference between processes is asymmetric — A can degrade B without B degrading A — which is worth knowing before designing anyone's day. ↩
-
Sophie Leroy, "Why is it so Hard to do My Work? The Challenge of Attention Residue when Switching Between Work Tasks," OBHDP (2009). The Zeigarnik effect is the older cousin here and I'd cite it carefully: the claim that you remember unfinished tasks better replicates poorly, while the claim that they intrude on attention until discharged holds up well. ↩
-
Masicampo and Baumeister (2011), converging with Leroy's later work. Worth sitting with: the plan does not have to be executed, or even executable soon. It has to be specific. ↩
-
Inzlicht and Schmeichel, "What Is Ego Depletion? Toward a Mechanistic Revision of the Resource Model of Self-Control," Perspectives on Psychological Science (2012). ↩
-
One evening, recalled, uninstrumented. I'm using it because it illustrates a documented mechanism cleanly, not because it demonstrates one. If the only support for this section were my evening, the section shouldn't exist. ↩
-
Hagger et al. (2016) and Vohs et al. (2021) are the two large preregistered nulls; Dang et al. (2025) is the intensity-based paradigm that recovers the effect. Note that ego depletion is about self-control rather than attention as such — the connection to attention capacity is an analogy commentators have drawn, not a test anyone ran. I'm using the history as a parable about mechanisms, not as direct evidence about attention. ↩
-
Baumeister, André, Southwick and Tice, "Self-control and limited willpower: Current status of ego depletion theory and research," Current Opinion in Psychology (2024) — the initial theory, in its own author's summary, "has been refined to emphasize conservation rather than resource exhaustion." The conservation finding is older than the replication crisis that forced it: Muraven, Shmueli and Burkley (2006) found that depleted participants who expected to need self-control later performed worse now. That is a budget, not a battery. ↩
-
With an obvious asterisk on the word "invariance," which holds for code and does not hold for agents. Agents absolutely do drop steps and then report that they didn't — I've written up an instance → where one told me a fix had shipped three times running while continuing to serve the placeholder. Whether that puts a clock on this entire argument is a separate essay, and I think it's the interesting one. ↩
-
Parnin and Rugaber, "Resumption strategies for interrupted programming tasks," Software Quality Journal (2011), drawn from roughly 10,000 recorded sessions across 86 programmers. ↩
-
Gloria Mark, "The Cost of Interrupted Work: More Speed and Stress" (CHI 2008). The figure is 23 minutes 15 seconds, and it describes the interval before returning to the original task, not the cost of resuming it. ↩
-
The interruption cadences are estimates of a typical hour rather than logged data, and I'm rounding hard. The two states are also a month apart, so the tooling is not the only thing that changed between them. The claim I'd defend is the ratio between the two states and the direction of the utilization figure, not the second decimal place of either. ↩