I was in middle school and stuck on World's Longest Platformer, a hundred-stage wall of a Scratch project that was extremely popular and comfortably beyond me. Then I noticed something: the whole thing was open in the editor, and in the editor I could just grab my own sprite and drag it past whatever was killing me. That worked. It worked every single time. And it was miserable, because I had to do it by hand.
So I went hunting through Scratch's block categories — Motion, go to, mouse pointer — and wired a keypress to teleport my sprite to wherever my cursor happened to be. One key, and every obstacle in the game became optional.1
I want to be precise about what happened there, because it is the same thing that has happened at every job I've had since. I didn't build the game. I didn't invent the cheat either — dragging the sprite past the obstacle was already the cheat, and it already worked. What I built was the thing that took the repetition out of a cheat I already had.
That's the whole move. I've been running it for over a decade and it has never once been sophisticated.
The résumé version of this essay is a boring list: at every position after my first, I automated something nobody asked me to automate. Fine. Lists don't argue. The interesting part is that not one of them was a decision — I have never sat down and resolved to be a person who automates things. Each time, I ran the same small calculation, and the calculation kept coming back yes.
Which is why I think you're probably not automating enough. Not because automation is virtuous; it isn't. Because the calculation is short, most people never run it, and the inputs it depends on have moved a great deal in the last two years.
The calculation
Three terms.
- Are you improving from a low floor, or maintaining a high one?
- Is doing it manually slower, or is correcting the automation's overreach slower?
- What risk attaches to each side?
That's all of it. Run those and the bounds fall out on their own — including every sober-sounding rule people mistake for ethics. "Warn, don't block." "A human retains merge authority." "The tool flags it; the tool never grades it." I hold all three and not one of them is a principle. They're what the arithmetic said, in a specific situation, with specific numbers.
If that sounds familiar, it's because I've made this argument one level down. Good code is code that saves you time → — every property we praise code for is a proxy for somebody's time, and most style arguments get clearer the moment you ask whose. This is the same move pointed at tools instead of code, except that the currency isn't quite time.
It's attention, and the two come apart more than you'd expect. Attention doesn't bank. It isn't one pool. It bills at wildly different rates for the same hour, and a task you've opened keeps charging you until it closes. The companion to this essay → works all of that out properly — the research, the mechanisms, what follows. This one only needs you to grant that hours are the wrong denomination, because every number below is wrong if you don't.
Here's where I first ran it.
In the summer of 2024 I was an executive assistant at Anytime AI, a legal-AI startup, and part of the job was turning event lead lists into CRM records. Twenty events, roughly a hundred leads each. Every export came in a different format, and the data underneath was human: the same person spelled three ways, work and personal emails for the same lawyer, entries duplicated across two events. For each row I had to work out who the person actually was, then go find them — usually the law office's website, occasionally a bar association registry. Amortized to a single legal practice, about half a minute per record.
Do that math. Two thousand records at thirty seconds is just under seventeen hours — two working days. Enough to argue about, not enough to settle anything, and over a summer it's the kind of number you eat.
I did about thirty of them and knew I wasn't going to make it.
The half-minute is amortized, but not to a large enough scale, and that's the trap. The actual work was: read a name, decide which of three spellings is real, open a registry, search, fail, open the firm's site, find the attorney page, pull five fields, then find your place again in a spreadsheet formatted differently from the last one. Nobody sustains that for more than about thirty minutes. If it were your only task you'd be genuinely on it maybe a quarter of the time — which turns seventeen hours of clock into sixty-eight-plus hours of day, and something like a week and a half of hell.
Seventeen hours was never the real number. The estimate was wrong by a factor of four, and wrong in a specific way: it priced a rate nobody can hold. Same task, same records, same thirty seconds — the only thing that changed is the unit.
If you price your own toil in hours, you will systematically under-automate. That's the first reason I think you're doing it.
The same question, two floors
The first term does most of the work, and it splits into two situations that produce completely opposite behavior.
When the floor is already dirty
So I wrote the pipeline. Python and BeautifulSoup, POSTing to bar registries and scraping web results, normalizing twenty inconsistent exports into one CSV shape and routing records per jurisdiction.
Three decisions in there I'd still defend. I narrowed the scope on purpose — the company only cared about Florida, New Jersey, Pennsylvania, and New York, all four of which had usable registries, and I made no attempt at general coverage. I let it automate a judgment, which I want to be honest about: deciding that a garbled name and a personal email belong to a particular attorney is a judgment, and the pipeline made it two thousand times without me. What made that defensible wasn't restraint, it was the third term — the judgment was checked against an external authority rather than invented, so every field came out as correct as the bar has it, which is a far weaker claim than "correct" and a far more defensible one. And I throttled it to one record per second so I wouldn't accidentally DoS a state bar association, which cost nothing — the run took half an hour instead of two minutes, and half an hour is free when you aren't the one sitting there.
The reason I could be that aggressive is the floor. The existing data was a mess, so I wasn't protecting quality, I was manufacturing it out of nothing — which makes available the strongest safety argument I've ever had: I don't think a single row came out worse than they already had it. When the floor is dirty, your downside is bounded by the floor. You cannot lose.
Then it grew. Deduplication, entry merging, more format handlers, tacked on as I hit the cases. Build and accretion together came to about a week. That is how you should build your own tools, and much more so now that iterating is cheap.
I told the COO I was writing a script, then showed up with it done. He liked it a great deal, mostly because the CRM migration it fed landed two weeks early. It still died with me: I zipped it up with documentation, but it was customized per sheet, and once the company moved to Zoho it stopped mattering.
When the floor is a working product
Two years later, at Chewy, the floor was a live interface four thousand customer care agents used to do their jobs, and my assignment was to pull a UI framework out from underneath it. React 17, Next 12, Material UI v4 and v5 both installed, an internal component library, CSS Modules, global CSS, custom wrappers. There was no canonical dropdown and no canonical heading — five controls that looked identical on screen were five implementations with five override histories, some resolving differently after client hydration than they had on the server.
Then the regression that rearranged how I think about all of this. A replacement selector didn't apply everywhere it needed to, so the newer library's typography defaults won in a handful of places and headings came out enormous. Lint passed. Types passed. Unit tests passed. The build was green. The interface was visibly broken.
Compilation does not prove visual correctness. Once you've watched that sentence be true, the arithmetic changes shape. On a dirty floor, automation manufactures quality. On a high floor it can only ever spend it — so the thing you build stops being a machine that does work and becomes a machine that produces evidence.
Which is what it was: Playwright driving Chromium across representative routes and isolated component states, against 102 mocked surfaces so the data couldn't move underneath a capture, compared against a baseline pinned to an exact commit, with semantic pixel diffs and cache fingerprints so a stale screenshot could never quietly promote itself into a baseline.2
Note what it doesn't do. It never merged anything. The delivery workflow I built on top of it — oneprompt, a 119-line skill that took one ticket at a time and scheduled it, pinned a baseline, isolated a branch, resolved it, ran the repository's checks, published, and followed up on the CI results — published draft only, every time. My own note from that summer says it better than I can now: I automated attention, not responsibility.
That workflow is also where the attention framing stopped being a metaphor for me. I measured it expecting speed and got 24.8 hours to merge before, 23.4 after — noise.3 By the stopwatch it did nothing whatsoever. What it did was make the same elapsed time cost less of me: each of those eight steps carried a small obligation to remember it, and the sequence stayed open in my head until the last one closed, billing continuously. Encoding the loop closed it, which is why five branches took first commits inside a 34-minute window. The attention it returned had to go somewhere.
A tool that saves zero time and real attention is a good tool. There is no line on any dashboard where that shows up.
And here's the boundary people skip: it wasn't enough. Structured manual testing — old build beside new build, a month of it — kept surfacing states the fixtures had never represented. Hover behavior. Modal positioning. Form state after a cancel. The tool narrowed the search space and then a person had to look. What I'd defend as craft there isn't the harness; it's that findings got classified before they got counted — migration-caused, accepted difference, pre-existing, environmental, uncertain. An undifferentiated bug list is a pile of work. A classified one is a decision.
I built the same instrument three times
I want to point at something I did not notice until somebody asked me about it directly.
At Anytime AI the product had four pillars: legal research, document analysis, document drafting, and general Q&A. Two of them — research and drafting — were simply dead in the lower environments, because the services they leaned on weren't connected down there. Not broken. Invisible. You could put a defect into that part of the UI and nothing would tell you until it hit staging, by which point the startup calculus had usually shipped it. That's a defensible trade when runway is the binding constraint; it also meant two of four pillars had no early warning at all. So I built the environment that didn't exist, mocking the API responses the frontend expected, and tested against my own fake backend. It surfaced some genuinely egregious things that thin testing was never going to catch on its own.
At NExT Consulting, I mocked the client's Quickbase so we could build purchase-order check-ins — including check-in-from-a-spreadsheet — without depending on their live instance.
At Chewy, 102 mocked surfaces under a pixel harness.
Three unrelated domains, three years, the same instrument, and it never once felt like a pattern while I was inside it. My honest account of why is unflattering: I'm a lazy developer, I notice the thing slowing me down, and I fix it so I can keep moving.
But look at what those three have in common, because it isn't work. Not one of those mocks did any part of my job for me. What they removed was blindness.
Which is where they connect back to the calculation, and it took me embarrassingly long to notice. An instrument doesn't automate the work. It makes the work priceable.
So when you can't see the failure, the first thing worth building isn't the script that does the task. It's the thing that makes the task's failure visible — because until that exists you aren't running the calculation at all. You're guessing, and then calling the guess a decision.
Why you're probably under-automating
The receipts are on the table, so here's the argument.
You're pricing it in hours. The delivery workflow above moved the median 24.8 hours to 23.4 — noise, indistinguishable from nothing, and it was one of the most useful things I built that summer. Run that through a time estimate and you don't build it. Run it through an attention one and it's obvious.
You're waiting to be told. Not one of these was assigned. At Anytime AI I mentioned the script and then arrived with it finished, which worked partly because the company hadn't built an intern playbook yet. A set of inventory soundness checks at NExT I derived myself, from watching how stock entered, left, and expired — then raised, and got confirmed almost immediately, because the client had nearly been burned by exactly that, twice. Nobody was ever going to file that ticket. The people who could see the problem were not the people who could see the fix.
You think automating means replacing judgment, which sounds dangerous, so you never start. But where judgment sits is also just the calculation, and run honestly it usually says leave it where it is — which is why the scary version of automation is also the rare one. The test is whether anything outside the tool can rule on the judgment it would be making, and this essay has a case on each side of it. The scary version of automation is rare because the math rarely licenses it, not because anyone forbade it.
You ran the calculation once, and it was a long time ago. All three terms have moved, and the cost of iterating on a tool has collapsed harder than the cost of building one.
And xkcd 1205 ↗ needs recalibrating. I cited that comic approvingly in the good-code essay and still use it. But be clear about its axes: it prices the time you spend against the time you save, and it was drawn when the expensive input was a developer typing. Both premises moved. The chart still works; the numbers you feed it shouldn't be hours of typing anymore.
I've been on the wrong side of that last pair myself. Back at Chewy, the screenshot suite got slow enough to be annoying as coverage grew, so I built sharded and multithreaded capture, ran several builds concurrently, and benchmarked the complete workflow rather than worker utilization. About seven percent. Compilation had been the dominant cost the entire time, and running multiple builds at once mostly just cooked my machine, so I deleted the scheduler.
The honest verdict isn't that sharding was a bad idea. It's that I ran the experiment against the wrong compiler. The whole result hinges on compilation dominating, and that is exactly the term a build-tooling upgrade moves — I'd run it again on the other side of a webpack-to-turbopack migration with the packages to match, and I'd expect a different number. A technically interesting optimization that doesn't move the bottleneck is a hobby. Checking whether the bottleneck moved is not.
Where the line falls
None of that is a case for automating everything. The calculation draws a line far more often than it hands over a whole task, and the line lands somewhere different every time. Twice below I built the thing and it still stopped well short of the work — once through the middle of the task, once around the authority to act on it.
Two oracles in one rubric
In the fall of 2025 I was a lead TA at Northeastern for a course with about 450 students and twenty TAs, each of us grading around twenty submissions a cycle. Style was part of the rubric: a minor infraction needed a comment, a major one cost a point.
That sounds like one task. It was two, and they had different oracles.
Finding an infraction was mechanically checkable — line 31 either exceeded the limit or it didn't, and pylint was right about it more reliably than I was. Deciding whether the infraction was minor or major was a judgment about a student's intent and the assignment's context, and there was no registry to POST to for that. Same term as the bar pipeline, opposite answer: the pipeline got to make a judgment two thousand times because an external authority could rule on it, and the aggregator got to make none because nothing could.
The second term priced it the same way. An auto-assigned deduction that came out wrong had to be found and reversed, and reversing one cost more than making it by hand in the first place. A missed infraction cost nothing extra, because the TA was reading the code anyway.
So the line ran through the middle of the task, at the seam between the half with an oracle and the half without.
Pylint was already wired into the autograder, dutifully flagging everything, and "everything" was the problem: one to two hundred lines per submission, most of it irrelevant to our schema, some of it flagging infractions that were genuinely necessary, and some of it flagging code that had shipped non-compliant in the handout itself. Files came back jumbled, warnings weren't grouped. Grading meant scroll the wall, spot a real violation, note the line, navigate to it, mark it in the platform, come back, find where you'd been, keep scrolling.
The tool I wrote did almost nothing. It read pylint's output and reorganized it — by file, then by category, then just the line numbers underneath. line-too-long: 14, 31, 32, 78. That was all a TA actually needed, and it fit on one screen. It never assigned a point. What went away was the place-keeping — half an hour per TA per assignment, fourteen TAs who actually ran it, ten assignments. Seventy-odd hours a semester, and not one of them was work. Cursor movement.
Most of the TAs in the org ended up running it.4 Of everything in this essay it automated the least, and more people used it than anything else I've built.
A low floor that licensed nothing
Every other time the calculation drew a line, the floor was high and that was the tell: the tool couldn't be allowed to spend quality it hadn't made. dupe-audit was the one where the first term said the opposite and the line landed in the same place anyway.
The floor was low. There was duplication everywhere — enough that any halfway competent detector found something on its first run. By the logic of the bar pipeline that licensed aggression, since a mess doesn't get worse for being cleaned.
Except the mess worked. Duplicated code is low-quality by every metric available, and it was also, right then, in production, passing tests, serving customers. The floor was low on the dimension I'd be improving and high on the dimension I'd be risking, and those weren't the same dimension.
Both directions of error priced out lopsided. A detector tuned to catch everything flagged everything — near-misses, coincidental shape, three functions that happened to walk a list the same way. Past some threshold the output stopped being a list of findings and became a wall, and a wall is the kind of problem no amount of concentration solves. It wouldn't have been abandoned for being wrong. It would have been abandoned because reading it cost more than ignoring it.
Under-detection cost almost nothing. Fewer proposals, and given how much slop was actually in there, fewer was still plenty. A conservative duplicate auditor was never going to run out of things to say.
Too eager cost the tool. Too timid cost nearly nothing. The second term pointed the opposite way from the first, and it pointed harder.
The third term settled the rest. Structural similarity isn't intent — two functions can be identical and mean different things, one a coincidence that will diverge next quarter and the other a genuine duplication, with nothing in the source to separate them. No external authority could rule on that, which is the test the aggregator failed and the bar pipeline passed.
So it proposes and it doesn't delete. It will argue for a consolidation all day and has no ability to perform one. It was in beta at 148 tests when I left.
Those are three of the rules I quoted at the top of this essay, and they were always the same rule with different numbers in it. The aggregator flags and never grades. oneprompt publishes and never merges. dupe-audit proposes and never deletes. The inventory checks warn and never block. Four implementations, and not one of them started as a principle.
Even failing is fine
NExT was the first place I had Claude Code, and I ended up running several sessions at once. The reason was not ambition. While one was working I had spare attention — genuinely idle capacity, expiring in real time, with nothing else to spend it on. An attempt built out of attention that was going to evaporate anyway does not cost what an attempt used to cost, and that changes which attempts are worth making.
Months later at Chewy, that's what Migration Cleanup Autopilot came out of. It was 120 files and 29,639 lines: a control plane for bounded migration work, with sealed task authority, disposable worktrees, repository-owned validation the agent couldn't rewrite, SQLite lifecycle evidence, bounded repair cycles, draft-only publication, and fail-closed behavior whenever the evidence wasn't good enough. Architecturally I still think it's right.
It has never been properly stood up — not because the idea didn't survive contact, but because standing it up needed infrastructure work I no longer had the runway for, and handing it to somebody else would have cost them more attention than it was going to return that quarter.5 It saved zero time. It had zero users. That's a failure and I'm not going to dress it up, but I want to be exact about which kind: it failed on circumstance, not on nature. Another month and I think I'd have had it running. If it worked, I'd put what it saves at multiple months of developer time a year.
Priced honestly, it cost about five hours of attention that had no alternative use, plus a token bill. And the failure wasn't sterile: the authority model it forced me to make explicit went straight into the handoff documentation I left behind, and into dupe-audit, which is built on exactly that separation.
There is a fundamental imbalance there, and it's new. A speculative project that fails outright, built with attention that was going to evaporate anyway, costs approximately nothing and can still throw off learnings that land somewhere real. When the marginal cost of an attempt approaches zero, the expected-value math on speculative automation inverts: things with a high probability of failure become rational to try. That was not true in 2024. I could not have afforded to be wrong at 29,639 lines.
So: a failure worth having, and I'd do it again — the same conclusion I reached about every unfinishable project I built as a teenager →, for entirely different reasons.
So, you
I don't think automation is a virtue, and I'm suspicious of people who do. It isn't discipline, it isn't seniority, and it certainly isn't a personality. It's a short calculation with three terms, and I've run it every time I've caught myself doing something too many times, for a reason that has nothing to do with productivity: I never want to be bored.
That needs one clarification, because it inverts how boredom usually gets described. Boredom isn't having too much to do, and it isn't having too little. You can spend attention all day and be bored out of your mind — thirty seconds a record, two thousand records, fully occupied and dying. You can also be flat out for two hours on something hard and not be bored for a second of it. Being bored out of your mind is what happens when you have attention free and nowhere worth putting it.
So automation was never how I do less. It's how I get attention back out of the tasks that were charging me for nothing, so I can put it somewhere that isn't boring. The laziness I copped to a few sections ago is real, but it's the mechanism, not the motive. The motive is that I would much rather be overloaded than idle.
So run the calculation, and run it properly. Price it in attention, not hours. Assume nobody is going to ask you to. Build the instrument before you build the machine. Move judgment when the third term says you can and leave it alone when it says you can't — the point is that you ran the term, not that you flinched. Then notice that the cost of being wrong has fallen through the floor, and take a swing at something that probably won't work.
A decade ago the swing was a keypress that moved a sprite to my cursor. It is still the same swing. The only thing that's changed is the exchange rate, and it's still moving.
Receipts: the ledger
Every automation I've run this calculation on, and what it said, roughly in the order I ran it. The second column prices both directions of being wrong — what it costs when the tool does too much, and what it costs when it misses.
| Automation | Floor | Overzealous vs. missing | Risk | Line |
|---|---|---|---|---|
| Scratch teleport-to-cursor | Basement — I kept dying | Both failure modes cost one keypress | None, it's a game | Automate all of it |
| GuardBot control app | — | — | — | Before the calculation |
| Bar-registry pipeline | Basement — dirty data | Bad matches checkable; misses keep the floor | Bounded by the floor; no row got worse | Automate the judgment |
| Anytime AI mock | The abyss — anything beats impossibility | More mock data is better; gaps ship bugs to prod | Bounded — the alternative was blindness | Build the instrument first |
| TA style aggregator | High — grading was already right | Correcting a deduction costs more than grading | No oracle can rule on intent | Reading only, never grading |
| Quickbase mock | The abyss — anything beats impossibility | More mock data is better; gaps ship bugs to stage | Bounded — the alternative was blindness | Build the instrument first |
| Inventory soundness checks | High — a live client system | False warning: a glance. False block: the user. | A block takes the detection with it | Around authority, not work |
| Alembic migration | High — live production data | At n = 1, false alarms are indistinguishable | An unverified checker on live data | Before the tool |
| Playwright pixel harness | High — a UI 4,000 agents use | Extra captures are free; a miss is a rollback | A green build proves nothing about pixels | Evidence, never action |
oneprompt | High — a working repo | Too eager burns tokens; too timid burns attention | Capped by draft-only publication | Draft only, every time |
| LunchBot | Basement — manual polling | False ping: twenty breaks. Miss: back to polling. | You can't un-interrupt twenty people | Automate; tune both error rates |
| MCA | High — a working repo | Bad drafts are unfixable; timid still clears chaff | Agent authority over a live repo | Sealed authority, fail-closed |
dupe-audit | Low — plenty of slop, but it all works | Over-flagging dilutes; a miss still leaves plenty | A wrong merge deletes working code | Propose, never delete |
Two got picked up by other people. One shipped to a client, and one internal mock dies and gets revived with every new engagement. Three died with me — the bar pipeline, oneprompt, and MCA. Two I never built, one I deleted after measuring it, and one failed at 29,639 lines and paid for itself anyway.
I'd run all of them again. The sharding too — just on the other side of the compiler change that made it pointless the first time.
Notes
-
From there it was infinite ammo in a shooter, air jumps in a platformer, whatever a given game was withholding from me. None of it was hard. That whole era is covered in how I learned to code →; the relevant part here is only that my first program automated a manual process rather than building anything. ↩
-
For scale on the migration itself: fifty-one pull requests opened, forty-six merged, and a late audit that found 166 production files still importing the legacy styling API after I thought I was finished. That's a different essay. ↩
-
The concurrency figure is the one I'd defend; I can't prove the workflow caused it rather than the deadline did. I'd still rather report a flat median than quote the number that flatters me. ↩
-
A new assignment would occasionally surface a lint category my grouping didn't handle. I graded much earlier than most people, so a patch usually went out the same day grading opened. That is less admirable than it sounds: I was the canary because I was impatient, not because I planned to be. ↩
-
The setup cost is what I'd watch if I built this again. A control plane that needs more standing-up than the work it absorbs in the time you actually have is a control plane you won't finish, however good the architecture is — and the term that decides it is your remaining runway, which is not a term I had written down anywhere. ↩